Hidayat, Muhammad Rifqi (2026) Pengembangan Large Language Model Untuk Root Cause Failure Analysis Pada Pembangkit Listrik Tenaga Uap Di Indonesia. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6026242018-Master_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (1MB) | Request a copy |
Abstract
Root Cause Failure Analysis (RCFA) merupakan pendekatan penting dalam industri Pembangkit Listrik Tenaga Uap (PLTU) untuk mengidentifikasi akar penyebab kegagalan peralatan atau sistem yang ada pada pembangkit. Namun, proses ini masih dilakukan secara manual dan bergantung pada keahlian insinyur berpengalaman. Penelitian ini bertujuan mengembangkan Large Language Model (LLM) dengan kemampuan domain spesifik RCFA PLTU di Indonesia melalui pendekatan fine-tuning berbasis Parameter-Efficient Fine-Tuning (PEFT) dengan metode Low-Rank Adaptation (LoRA). Dataset penelitian dibangun dari 15 laporan historis RCFA PLTU di Indonesia periode 2022-2024 yang bersifat naratif dan tidak terstruktur, diekstraksi menjadi format JSON terstruktur, divalidasi, kemudian ditransformasikan menjadi 322 pasangan Question-Answer (QA), termasuk augmentasi berupa parafrase pertanyaan. Fine-tuning dilakukan menggunakan pendekatan QLoRA (kuantisasi 4-bit dengan adapter LoRA) pada tiga LLM open-source berukuran kecil, yaitu Qwen3.5-4B, Gemma3-4B, dan Llama-3.2-3B-Instruct dengan konfigurasi hyperparameter LoRA yang identik pada ketiga model, pada lingkungan komputasi terbatas yaitu GPU 12GB VRAM. Hasil evaluasi menunjukkan bahwa ketiga model berhasil mempelajari pengetahuan domain RCFA yang sebelumnya tidak dimiliki oleh base model dengan tingkat keberhasilan uji memorisasi dengan correct sebesar 92,55% pada Qwen3.5-4B, 91,61% pada Gemma3-4B, dan 86,65% pada Llama-3.2-3B. Pengujian generalisasi terhadap pertanyaan yang diparafrase juga menunjukkan Qwen3.5-4B mencapai keseimbangan terbaik dengan tingkat correct sebesar 90% tanpa satupun jawaban incorrect, sedangkan Gemma3-4B mencapai akurasi correct yang setara tetapi dengan tingkat kesalahan incorrect yang lebih tinggi, dan Llama-3.2-3B menunjukkan performa terendah tetapi tetap mengalami peningkatan dibandingkan base model. Penelitian ini menunjukkan bahwa pembuatan dataset domain spesifik RCFA PLTU dan pengembangan LLM dengan kemampuan domain tersebut berhasil dicapai melalui pendekatan PEFT LoRA yang efektif digunakan pada lingkungan komputasi terbatas hanya dengan laporan RCFA yang tidak terstruktur.
==================================================================================================================================
Root Cause Failure Analysis (RCFA) is an important approach in the coal-fired steam power plant (PLTU) industry for identifying the root causes of equipment or system failures within a power plant. However, this process is still carried out manually and relies heavily on the expertise of experienced engineers. This research aims to develop a Large Language Model (LLM) with domain-specific capability for RCFA at coal-fired steam power plants in Indonesia through a fine-tuning approach based on Parameter-Efficient Fine-Tuning (PEFT) using the Low-Rank Adaptation (LoRA) method. The research dataset was constructed from 15 historical RCFA reports from an Indonesian coal-fired steam power plant covering the 2022-2024 period, which were narrative and unstructured in nature. These reports were extracted into a structured JSON format, validated, and then transformed into 322 Question-Answer (QA) pairs, including augmentation through question paraphrasing. Fine-tuning was performed using the QLoRA approach (4-bit quantization with LoRA adapters) on three small open-source LLMs, namely Qwen3.5-4B, Gemma3-4B, and Llama-3.2-3B-Instruct, using an identical LoRA hyperparameter configuration across all three models, within a resource-constrained computing environment using a 12GB VRAM GPU. The evaluation results show that all three models successfully learned RCFA domain knowledge that the base models did not previously possess, achieving memorization test success rates of 92.55% for Qwen3.5-4B, 91.61% for Gemma3-4B, and 86.65% for Llama-3.2-3B. Generalization testing on paraphrased questions also showed that Qwen3.5-4B achieved the best balance, with a correct rate of 90% and no incorrect answers at all, while Gemma3-4B achieved a comparable correct accuracy but with a higher incorrect rate, and Llama-3.2-3B showed the lowest performance yet still exhibited improvement over its base model. This research demonstrates that the construction of a domain-specific RCFA dataset for coal-fired steam power plants and the development of an LLM with such domain capability were successfully achieved through a PEFT LoRA approach that is effective for use in resource-constrained computing environments using only unstructured RCFA reports.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | Fine-tuning, Low-Rank Adaptation, Large Language Model, Pembangkit Listrik Tenaga Uap, Root Cause Failure Analysis, Fine-tuning, Low-Rank Adaptation, Large Language Model, Coal-fired Power Plants, Root Cause Failure |
| Subjects: | Q Science > QA Mathematics > QA76.76.E95 Expert systems Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science) T Technology > TJ Mechanical engineering and machinery > TJ164 Power plants--Design and construction |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 59101-(S2) Master Thesis |
| Depositing User: | Muhammad Rifqi Hidayat |
| Date Deposited: | 30 Jul 2026 08:23 |
| Last Modified: | 30 Jul 2026 08:23 |
| URI: | http://repository.its.ac.id/id/eprint/140157 |
Actions (login required)
![]() |
View Item |
