Pengembangan Sistem Text-to-Speech Berbasis StyleTTS2 dengan Pemodelan Interaksi Fonem Prosodi untuk Adaptasi Prosodi Bahasa Indonesia ke Bahasa Jawa

Arisdani, Eldintaro Farrandi (2026) Pengembangan Sistem Text-to-Speech Berbasis StyleTTS2 dengan Pemodelan Interaksi Fonem Prosodi untuk Adaptasi Prosodi Bahasa Indonesia ke Bahasa Jawa. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221299-Undergraduate_Thesis.pdf] Text
5025221299-Undergraduate_Thesis.pdf
Restricted to Repository staff only

Download (10MB) | Request a copy

Abstract

Sintesis ujaran lintas bahasa pada kondisi bahasa berdaya rendah ekstrem menghadapi hambatan Input Phonetic Mismatch akibat ketiadaan modul Grapheme-to-Phoneme (G2P) mandiri pada bahasa target. Penggunaan proksi G2P serta adanya keunikan prosodi bahasa sumber rentan memicu ketidakakuratan pelafalan dan kebocoran aksen yang mendegradasi kealamian prosodi bahasa target. Di sisi lain, metode full fine-tuning pada dataset yang sangat terbatas rentan memicu kelupaan katastropik catastrophic forgetting serta merusak stabilitas alinyemen model dasar.

Penelitian ini mengusulkan kerangka kerja Parameter-Efficient Fine-Tuning (PEFT) pada arsitektur StyleTTS2 untuk adaptasi lintas bahasa dari Bahasa Indonesia ke Bahasa Jawa menggunakan dataset monolingual berdurasi minimal 32 menit. Pendekatan ini melokalisasi pembaruan parameter pada dua modul tambahan kustom: Language Phoneme Embedding Processor (LPEP) pada tahap awal untuk menyelaraskan representasi fonem secara dinamis menggunakan skema weight cloning dan identity initialization, serta Phoneme Prosody Interaction Module (PPIM) pada tahap lanjutan yang memanfaatkan gated cross-modal attention untuk menyatukan teks dengan variabilitas prosodi secara temporal. Alur pembaruan parameter terisolasi ini dikendalikan secara bertahap oleh sirkuit 2-Stage PEFT Controller.

Hasil pengujian subjektif menunjukkan peningkatan kealamian dengan skor MOS-N sebesar 4,328 (mengungguli full fine-tuning sebesar 4,150), CMOS-N bersih sebesar +0,860, dan akurasi pelafalan MOS-PA sebesar 4,287. Secara objektif, studi ablasi membuktikan integrasi LPEP dan PPIM memicu efek sinergi manifol laten yang mereduksi Word Error Rate (WER) model secara drastis menjadi 0,2552, dari sebelumnya 0,5506 (PPIM saja) dan 0,3758 (LPEP saja). Analisis klasterisasi t-SNE global mengonfirmasi keberhasilan isolasi gaya wicara (style separation) dengan Silhouette Score positif sebesar 0,862, serta efisiensi komputasi Real-Time Factor (RTF) sebesar 0,0276. Integrasi LPEP dan PPIM terbukti menjadi solusi arsitektur PEFT yang efisien, stabil, dan unggul untuk preservasi bahasa daerah berkategori low-resource.
============================================================
Cross-lingual speech synthesis under extreme low-resource conditions faces the Input Phonetic Mismatch obstacle due to the absence of a standalone Grapheme-to-Phoneme (G2P) module for the target language. The use of a proxy G2P and the presence of unique source language prosody tend to trigger pronunciation inaccuracies and accent leakage, which degrade the naturalness of the target language prosody. On the other hand, full fine-tuning on extremely limited datasets is prone to triggering catastrophic forgetting and disrupting the alignment stability of the base model.
=================================================================================================================================
This study proposes a Parameter-Efficient Fine-Tuning (PEFT) framework on the StyleTTS2 architecture for cross-lingual adaptation from Indonesian to Javanese using a monolingual dataset with a minimal duration of 32 minutes. This approach localizes parameter updates to two novel custom modules: the Language Phoneme Embedding Processor (LPEP) at the upstream stage to dynamically align phoneme representations using weight cloning and identity initialization schemes, and the Phoneme Prosody Interaction Module (PPIM) at the downstream stage that utilizes gated cross-modal attention to temporally unify text with prosodic variability. This isolated parameter update workflow is gradually regulated by a 2-Stage PEFT Controller circuit.

Subjective evaluation results demonstrate improved naturalness with a MOS-N score of 4.328 (outperforming full fine-tuning at 4.150), a net CMOS-N of +0.860, and a pronunciation accuracy MOS-PA of 4.287. Objectively, ablation studies prove that the integration of LPEP and PPIM triggers a latent manifold synergistic effect that drastically reduces the model's Word Error Rate (WER) to 0.2552, down from 0.5506 (PPIM only) and 0.3758 (LPEP only). Global t-SNE clustering analysis confirms the success of style separation with a high positive Silhouette Score of 0.862, alongside a computational efficiency with a Real-Time Factor (RTF) of 0.0276. The integration of LPEP and PPIM proves to be an efficient, stable, and superior PEFT architectural solution for low-resource regional language preservation.

Item Type: Thesis (Other)
Uncontrolled Keywords: Text-to-Speech, Cross-Lingual, Bahasa Jawa, Parameter-Efficient Fine-Tuning, StyleTTS2, LPEP, PPIM, Javanese Language.
Subjects: Q Science > QA Mathematics > QA336 Artificial Intelligence
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering
Depositing User: Eldintaro Farrandi
Date Deposited: 31 Jul 2026 07:18
Last Modified: 31 Jul 2026 07:18
URI: http://repository.its.ac.id/id/eprint/140808

Actions (login required)

View Item View Item