Anggrayni, Elok (2016) Peningkatan Kualitas Sintesis Ucapan (Natural Speech Synthesis) dengan Pendekatan Klasifikasi Voiced/Unvoiced di Lingkungan Bising Berbasis Estimasi Instantaneous Frequency Amplitude Spectrum (IFAS). Doctoral thesis, Institut Teknologi Sepuluh Nopember.
|
Text
02311960010001-Doctoral.pdf - Accepted Version Restricted to Repository staff only Download (6MB) | Request a copy |
Abstract
Sintesis ucapan merupakan teknologi pembangkitan ucapan buatan dari teks yang natural dan mudah dipahami serta berpotensi diterapkan pada bidang kesehatan, pendidikan, dan teknologi bantu bagi penyandang disabilitas. Kualitas sintesis ucapan sangat dipengaruhi oleh ketepatan estimasi fitur akustik, khususnya frekuensi fundamental (F0). Penelitian disertasi ini bertujuan mengembangkan metode estimasi F0 berbasis Instantaneous Frequency Amplitude Spectrum (IFAS) dengan Gaussian Window untuk meningkatkan representasi struktur harmonik dan kualitas sintesis ucapan Bahasa Indonesia. Metode yang diusulkan memanfaatkan IFAS sebagai turunan Short-Time Fourier Transform (STFT) dengan pendekatan bandpass filter bank, serta klasifikasi voiced/unvoiced berbasis harmonicity measure. Evaluasi performansi IFAS dilakukan menggunakan basis data ucapan Bahasa Indonesia dalam lingkungan bersih dan bising. Hasil estimasi F0 selanjutnya diimplementasikan pada sistem sintesis ucapan berbasis Hidden Markov Model (HMM), Deep Neural Network (DNN), dan Tacotron2. Kinerja sistem dievaluasi menggunakan pengukuran objektif dan subjektif. Hasil pengujian subjektif menunjukkan bahwa metode yang diusulkan menghasilkan nilai terbaik hingga 4,43/5,00 untuk pembicara laki-laki dan 4,07/5,00 untuk pembicara wanita pada berbagai ekspresi emosional. Hasil pengujian objektif menunjukkan bahwa Tacotron2 menghasilkan nilai terbaik sebesar 1,06–2,60 untuk pembicara laki-laki dan 1,31–2,25 untuk pembicara wanita. Hasil penelitian disertasi ini menunjukkan bahwa IFAS Gaussian Window mampu menghasilkan estimasi F0 yang lebih akurat dan lebih robust terhadap gangguan kebisingan dibandingkan metode cepstrum dan autokorelasi. Metode ini juga mengurangi diskontinuitas kurva IF sehingga menghasilkan representasi struktur harmonik yang lebih baik dan berkontribusi terhadap peningkatan kualitas sintesis ucapan Bahasa Indonesia.
=====================================================================================================================================
Speech synthesis is a technology for generating natural and easily understandable synthetic speech from text, with potential applications in the fields of healthcare, education, and assistive technology for people with disabilities. The quality of speech synthesis is greatly influenced by the accuracy of acoustic feature estimates, particularly the fundamental frequency (F0). This dissertation research aims to develop an F0 estimation method based on the Instantaneous Frequency Amplitude Spectrum (IFAS) with a Gaussian window to improve the representation of harmonic structure and the quality of Indonesian speech synthesis. The proposed method utilizes IFAS as a derivative of the Short-Time Fourier Transform (STFT) with a bandpass filter bank approach, as well as voiced/unvoiced classification based on a harmonicity measure. The performance of IFAS was evaluated using a database of Indonesian speech in both clean and noisy environment. The resulting F0 estimates were then implemented in speech synthesis systems based on the Hidden Markov Model (HMM), Deep Neural Network (DNN), and Tacotron2. System performance was evaluated using both objective and subjective metrics. The results of the subjective testing showed that the proposed method produced the best scores of up to 4.43/5.00 for male speakers and 4.07/5.00 for female speakers across various emotional expressions. The results of objective testing show that Tacotron2 produces the best values, ranging from 1.06 to 2.60 for male speakers and from 1.31 to 2.25 for female speakers. The results of this dissertation research show that the IFAS Gaussian Window can produce more accurate F0 estimates that are more robust against noise interference compared to the cepstrum and autocorrelation methods. This method also reduces discontinuities in the IF curve, resulting in a better representation of the harmonic structure and contributing to improved Indonesian speech synthesis quality.
| Item Type: | Thesis (Doctoral) |
|---|---|
| Uncontrolled Keywords: | Bahasa Indonesia, Deep Neural Network, Hidden Markov Model, instantaneous frequency, sintesis ucapan, Tacotron2 Indonesian, Deep Neural Network, Hidden Markov Model, instantaneous frequency, speech synthesis, Tacotron2 |
| Subjects: | Q Science Q Science > QC Physics Q Science > QC Physics > QC221 Acoustics. Sound T Technology > T Technology (General) T Technology > T Technology (General) > T57.5 Data Processing T Technology > T Technology (General) > T59.7 Human-machine systems. |
| Divisions: | Faculty of Industrial Technology and Systems Engineering (INDSYS) > Physics Engineering > 30001-(S3) PhD Thesis |
| Depositing User: | Elok Anggrayni |
| Date Deposited: | 05 Oct 2026 00:53 |
| Last Modified: | 05 Oct 2026 00:53 |
| URI: | http://repository.its.ac.id/id/eprint/145157 |
Actions (login required)
![]() |
View Item |
