Peningkatan Kinerja Automatic Speech Recognition Bahasa Daerah Indonesia Menggunakan Generative Fusion Decoding Termodifikasi

Santosa, Agung (2026) Peningkatan Kinerja Automatic Speech Recognition Bahasa Daerah Indonesia Menggunakan Generative Fusion Decoding Termodifikasi. Doctoral thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 7022212008-Doctoral.pdf] Text
7022212008-Doctoral.pdf
Restricted to Repository staff only

Download (6MB) | Request a copy

Abstract

Model pengenalan ujaran otomatis dan model bahasa pralatih berbobot terbuka memungkinkan pengembangan sistem speech AI tanpa melatih model besar dari awal. Namun, penerapannya untuk bahasa daerah Indonesia terkendala keterbatasan korpus ujaran berlabel, sementara pelatihan ulang model akustik menuntut ratusan jam ujaran bertranskripsi yang tidak tersedia. Disertasi ini meningkatkan kinerja pengenalan ujaran Bahasa Jawa dan Bahasa Sunda melalui Generative Fusion Decoding (GFD) termodifikasi, yaitu memadukan model akustik yang dibekukan dan model bahasa pada tahap dekoding tanpa memperbarui satu pun parameter model akustik. Dua alur metodologis dikembangkan dalam satu kerangka fusi yang sama. Jalur pertama, GFD-LogitCalib, memisahkan dua peran yang pada GFD asli dipikul satu parameter, yaitu penyetaraan magnitudo dan pembobotan relatif. Faktor kalibrasi α diturunkan dari distribusi rasio empirik magnitudo log-probabilitas, sedangkan r menyatakan bobot model bahasa. Pada korpus INDspeech07, tingkat galat kata turun dari 52,25% menjadi 44,83% untuk Bahasa Jawa dan dari 53,28% menjadi 47,88% untuk Bahasa Sunda, yaitu 14,2% dan 10,1% secara relatif. Rentang r yang menghasilkan perbaikan meluas sekitar tiga kali lipat, dan konfigurasi optimalnya bertransfer ke korpus buta OpenSLR tanpa penyetelan ulang. Jalur kedua, GFD-MambaByte, menghapus aproksimasi percabangan tokenisasi yang ditanggung setiap instansiasi GFD terdahulu dengan memakai model bahasa state-space tingkat byte—sepanjang pengetahuan penulis, instansiasi GFD yang pertama berbasis arsitektur tersebut—sehingga probabilitas hipotesis terfaktorisasi secara eksak. Adaptasi ke bahasa target dikerjakan semata-mata pada teks publik melalui continual pre-training LoRA tanpa satu pun ujaran berlabel, sedangkan cache trie-prefiks dengan ekstensi O(1) menjadikannya layak dijalankan pada satu GPU V100. Pada test set speaker-independent INDspeech_NEWSTRA_EthnicSR, tingkat galat kata turun 6,29 dan 6,92 poin persentase untuk Bahasa Jawa dan Bahasa Sunda, yaitu 10,8% dan 11,2% secara relatif, keduanya dengan p < 0,001. Kedua jalur merupakan komplemen tanpa pelatihan bagi model yang di-finetune secara terawasi, bukan penggantinya; angka absolutnya tidak diperbandingkan langsung karena korpus dan protokol evaluasinya berbeda. Sebuah contoh implementasi yang menerapkan prinsip yang sama pada lapisan berbeda, yaitu sistem perintah suara Bahasa Indonesia pada perangkat tepi, disajikan tersendiri di luar klaim kontribusi utama.
=====================================================================================================================================
Automatic speech recognition models and open-weight pretrained language models enable the development of speech AI systems without requiring large-scale model training from scratch. However, their application to Indonesian regional languages remains constrained by the scarcity of labeled speech corpora, while retraining acoustic models requires hundreds of hours of transcribed speech that are unavailable. This dissertation improves speech recognition for Javanese and Sundanese through a modified Generative Fusion Decoding (GFD) framework that combines a frozen acoustic model with a language model during decoding without updating any acoustic-model parameters. Two methodological approaches are developed within this unified fusion framework. The first approach, GFD-LogitCalib, separates the dual roles assigned to a single parameter in the original GFD, namely magnitude equalization and relative weighting. A calibration factor (α) is derived from the empirical distribution of log-probability magnitude ratios, while *r* represents the language model weight. On the INDspeech07 corpus, the Word Error Rate (WER) decreases from 52.25% to 44.83% for Javanese and from 53.28% to 47.88% for Sundanese, corresponding to relative improvements of 14.2% and 10.1%, respectively. Furthermore, the range of *r* values yielding performance improvements expands by approximately threefold, and the optimal configuration generalizes to the blind OpenSLR corpus without requiring additional tuning. The second approach, GFD-MambaByte, eliminates the tokenization-branching approximation inherent in previous GFD implementations by employing a byte-level state-space language model—the first known GFD implementation on this architecture—allowing hypothesis probabilities to factorize exactly. Adaptation to the target language is performed using only publicly available text through LoRA-based continual pretraining without any labeled speech data, while a prefix-trie cache with O(1) extension enables execution on a single NVIDIA V100 GPU. On the speaker-independent test set of the INDspeech_NEWSTRA_EthnicSR corpus, the WER decreases by 6.29 and 6.92 percentage points for Javanese and Sundanese, corresponding to relative improvements of 10.8% and 11.2%, respectively, with both results achieving statistical significance (*p* < 0.001). Both approaches serve as training-free complements to supervised fine-tuning rather than replacements, and their absolute performance values are not directly compared because the evaluation corpora and experimental protocols differ. As an additional implementation example, the same underlying principle is applied at a different system layer to develop an Indonesian voice-command system for edge devices; however, this application is presented separately and is not included as part of the dissertation's primary contribution.

Item Type: Thesis (Doctoral)
Uncontrolled Keywords: bahasa Jawa dan Sunda, fusi dekoding, kalibrasi logit, model bahasa tingkat byte, pengenalan ujaran otomatis, automatic speech recognition, byte-level language models, decoding fusion, Javanese and Sundanese, logit calibration.
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7882.S65 Automatic speech recognition.
T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7895.S65 Speech recognition systems
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Electrical Engineering > 20001-(S3) PhD Thesis
Depositing User: Agung Santosa
Date Deposited: 07 Aug 2026 03:36
Last Modified: 07 Aug 2026 03:36
URI: http://repository.its.ac.id/id/eprint/144201

Actions (login required)

View Item View Item