Cahyani, Fatma (2026) Normalisasi Bahasa Indonesia Informal Berbasis Large Language Model (LLM) Untuk Sistem Interaksi Robot Servis Kampus ITS. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6026241015-Master_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (7MB) | Request a copy |
Abstract
Sistem layanan informasi berbasis suara membutuhkan kemampuan untuk memahami pertanyaan pengguna yang sering disampaikan dalam bentuk informal, tidak baku, singkat, dan bercampur slang. Penelitian ini bertujuan mengembangkan sistem normalisasi Bahasa Indonesia informal menjadi formal dan semantic retrieval untuk layanan informasi Kampus Institut Teknologi Sepuluh Nopember (ITS). Dataset penelitian disusun dari dokumen resmi dan informasi pada situs resmi ITS yang mencakup 39 topik, dengan 1.500 pasangan pertanyaan–jawaban formal sebagai ground truth yang diperluas menjadi 7.500 variasi pertanyaan informal untuk pelatihan model. Metode yang digunakan meliputi fine-tuning model LLM Qwen3-14B dan Ministral 3 14B menggunakan SFT-LoRA, serta retrieval berbasis cosine similarity pada PostgreSQL dengan ekstensi pgvector. Sistem diimplementasikan sebagai pipeline end-to-end yang terdiri dari Speech-to-Text, normalisasi, embedding, retrieval, verbalisasi jawaban, dan Text-to-Speech, dengan mekanisme fallback untuk pertanyaan di luar cakupan basis data.
Hasil pengujian normalisasi menunjukkan Ministral 3 14B mengungguli Qwen3-14B dengan konfigurasi terbaik pada learning rate 2×10⁻⁴ dan LoRA r=16 yang memperoleh BLEU 0,7907, ROUGE-1 0,8981, ROUGE-L 0,8931, cosine similarity 0,8558, dan semantic similarity 0,9673. Dibandingkan dengan base model Ministral 3 14B secara zero-shot, proses fine-tuning meningkatkan semantic similarity dari 0,7532 menjadi 0,9673 atau sebesar 28,43%, mengonfirmasi bahwa adaptasi SFT-LoRA efektif untuk tugas normalisasi bahasa informal pada domain layanan informasi kampus. Pada pengujian retrieval terhadap 200 data uji, sistem memperoleh Hit@1 sebesar 0,9600, MRR sebesar 0,9690, dan answer accuracy sebesar 0,9400. Pengujian answerability (keputusan menjawab atau fallback) menghasilkan accuracy 0,9450, precision 0,9009, recall 1,000, dan F1-score 0,9479. Hasil ini menunjukkan bahwa normalisasi bahasa informal dapat mendukung peningkatan kualitas retrieval jawaban pada sistem layanan informasi kampus berbasis suara.
========================================================================================================================================
Voice-based information service systems require the ability to understand user queries that are frequently expressed in informal, non-standard, abbreviated, and slang-mixed forms. This research aims to develop an informal-to-formal Indonesian language normalization and semantic retrieval system for the information services of Institut Teknologi Sepuluh Nopember (ITS). The research dataset was constructed from official campus documents and information available on the ITS official website, covering 39 topics, comprising 1,500 formal question–answer pairs as ground truth that were further augmented into 7,500 informal question variations for model training. The methodology encompasses fine-tuning the Qwen3-14B and Ministral 3 14B large language models using SFT-LoRA, along with retrieval based on cosine similarity in PostgreSQL with the pgvector extension. The system was implemented as an end-to-end pipeline consisting of Speech-to-Text, normalization, embedding, retrieval, answer verbalization, and Text-to-Speech, incorporating a fallback mechanism for queries outside the scope of the ground truth database.
The normalization evaluation results demonstrate that Ministral 3 14B outperformed Qwen3-14B, with the best configuration at a learning rate of 2×10⁻⁴ and LoRA r=16 achieving a BLEU score of 0.7907, ROUGE-1 of 0.8981, ROUGE-L of 0.8931, cosine similarity of 0.8558, and semantic similarity of 0.9673. Compared to the zero-shot Ministral 3 14B base model, the fine-tuning process improved semantic similarity from 0.7532 to 0.9673, representing a 28.43% improvement, confirming that SFT-LoRA adaptation is effective for informal language normalization within the campus information service domain. In the retrieval evaluation on 200 test queries, the system achieved a Hit@1 of 0.9600, MRR of 0.9690, and top-1 accuracy of 0.9400. The answerability evaluation (answer-versus-fallback decision) yielded an accuracy of 0.9450, precision of 0.9009, recall of 1.0000, and F1-score of 0.9479. These results indicate that informal language normalization can meaningfully support improved answer retrieval quality in voice-based campus information service systems.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | Normalisasi Kalimat Informal, Semantic Retrieval, SFT-LoRA, Large Language Model Informal Sentence Normalization, Semantic Retrieval, SFT-LoRA, Large Language Model |
| Subjects: | T Technology > T Technology (General) > T58.5 Information technology. IT--Auditing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 59101-(S2) Master Thesis |
| Depositing User: | Fatma Cahyani |
| Date Deposited: | 27 Jul 2026 03:44 |
| Last Modified: | 27 Jul 2026 03:44 |
| URI: | http://repository.its.ac.id/id/eprint/137640 |
Actions (login required)
![]() |
View Item |
