Amoriza, Muhammad Nabil Akhtar Raya (2026) Sistem Multi-Agen Dengan Mekanisme Self-Reflection Untuk Pembangkitan Narasi Laporan Evaluasi Diri. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221021-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (3MB) | Request a copy |
Abstract
Akreditasi program studi bidang informatika dan komputer oleh LAM INFOKOM mewajibkan penyusunan dokumen Laporan Evaluasi Diri (LED). Proses tersebut melibatkan analisis yang mendalam, ekstraksi berbagai sumber informasi, serta membutuhkan waktu dan sumber daya yang tinggi. Upaya otomatisasi menggunakan Model Bahasa Besar secara langsung masih menghadapi tantangan berupa penurunan kualitas penalaran pada dokumen panjang, kebocoran data historis, dan halusinasi akibat kompleksitas konteks masukan. Untuk mengatasi permasalahan tersebut, penelitian ini mengusulkan sistem semi-otomatis berbasis Multi-Agent System (MAS) yang mengintegrasikan pendekatan Decomposed Information Retrieval (DecomposedIR) dan kerangka kerja Reflexion untuk menghasilkan draf narasi pada sembilan kriteria utama instrumen akreditasi LAM INFOKOM. Metodologi penelitian mengorkestrasikan empat agen spesialis (Ekstraktor, Analis, Penulis, dan Evaluator) beserta satu agen Agregator menggunakan kerangka kerja LangGraph. Proses pembangkitan mengikuti siklus PPEPP (Penetapan, Pelaksanaan, Evaluasi, Pengendalian, dan Peningkatan), dengan dekomposisi proses pembangkitan hingga tingkat sub-fase dan pemanfaatan template. Sistem menggunakan model bahasa besar regional Gemma-SEA-LION-v4-27B-IT sebagai model utama, sedangkan mekanisme Reflexion diterapkan melalui umpan balik iteratif dari agen Evaluator kepada agen Penulis untuk meningkatkan kualitas draf secara bertahap. Kinerja sistem dibandingkan dan dievaluasi terhadap sistem penelitian terdahulu melalui serangkaian pengujian eksperimen yang mencakup perbandingan arsitektur, sensitivitas dataset, ablasi model, ablasi strategi prompt, optimasi penugasan agen, dan pengujian pada dataset tambahan. Evaluasi dilakukan menggunakan metrik NLP (BLEU, ROUGE, BERTScore, dan BARTScore), metrik berbasis LLM-as-a-Judge (Faithfulness, Structural Adherence, dan Academic Tone), metrik operasional seperti Negative String Presence (NSP), Average Revision Count (ARC), serta kesiapan draf (R_draft dan R_manifest). Hasil menunjukkan bahwa konfigurasi ablasi model menghasilkan performa terbaik dengan nilai R_manifest sebesar 0,902, ARC sebesar 2,23, serta Faithfulness sebesar 0,824. Secara keseluruhan, pendekatan MAS mampu menghasilkan draf yang lebih konsisten, koheren, dan memiliki kesesuaian yang lebih tinggi terhadap struktur serta karakteristik narasi LED dibandingkan metode baseline. Meskipun demikian, sistem yang diusulkan diposisikan sebagai alat bantu penyusunan draf sehingga hasil masih memerlukan proses validasi dan penyuntingan oleh penyusun maupun pakar akreditasi sebelum digunakan dalam proses akreditasi secara nyata.
==================================================================================================================================
The accreditation process for informatics and computer science study programs conducted by LAM INFOKOM requires the preparation of a comprehensive Self Assessment Report (Laporan Evaluasi Diri, LED). Developing this document involves extensive analysis, extraction from multiple sources of information, and considerable time and human effort. Previous attempts to automate this process using Large Language Models (LLMs) through direct prompting have been challenged by degraded reasoning performance in long-document generation, historical data leakage, and hallucinations caused by increasingly complex input contexts. To address these limitations, this research proposes a semi-automated Multi-Agent System (MAS) that integrates the Decomposed Information Retrieval (DecomposedIR) approach with the Reflexion framework to generate draft narratives for the nine main criteria of the LAM INFOKOM accreditation instrument. The proposed system orchestrates four specialized agents (Extractor, Analyst, Writer, and Evaluator) together with an Aggregator agent using the LangGraph framework. Document generation follows the PPEPP cycle (Establishment, Implementation, Evaluation, Control, and Improvement), where the drafting process is decomposed into sub-phases and templates are utilized. The system employs Gemma-SEA-LION-v4-27B-IT as its primary large language model, while the Reflexion framework is implemented through an iterative feedback loop in which the Evaluator agent reviews drafts and provides structured feedback to the Writer agent for subsequent revisions. System performance is evaluated against a previous study's approach through six experimental scenarios covering architectural comparison, dataset sensitivity, model ablation, prompt strategy ablation, model assignment optimization, and evaluation on an additional dataset. The proposed system is evaluated using conventional natural language processing metrics (BLEU, ROUGE, BERTScore, and BARTScore), LLM-as-a-Judge metrics (Faithfulness, Structural Adherence, and Academic Tone), as well as operational metrics including Negative String Presence (NSP), Average Revision Count (ARC), and draft readiness (R_draft and R_manifest). Experimental results show that the model ablation strategy achieves the best overall performance, obtaining an R_manifest score of 0.902, an ARC of 2.23, and a Faithfulness score of 0.824. Overall, the proposed MAS produces drafts that are more consistent, coherent, and better aligned with the structural and narrative characteristics of LAM INFOKOM Self Assessment Reports than the baseline approach. Nevertheless, the proposed system is intended as a drafting assistance tool and does not replace human authors. Consequently, the generated drafts still require validation and revision by accreditation experts before being used in actual accreditation processes.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Laporan Evaluasi Diri, LLM-as-a-Judge, Model Bahasa Besar, Rekayasa Prompt, Self-Reflection, Sistem Multi-Agen. Large Language Model, LLM-as-a-Judge, Multi-Agent System, Prompt Engineering, Self Assessment Report, Self-Reflection. |
| Subjects: | Q Science > QA Mathematics > QA336 Artificial Intelligence T Technology > T Technology (General) > T57.5 Data Processing T Technology > T Technology (General) > T58.62 Decision support systems T Technology > T Technology (General) > T58.8 Productivity. Efficiency |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Muhammad Nabil Akhtar Raya Amoriza |
| Date Deposited: | 23 Jul 2026 02:40 |
| Last Modified: | 23 Jul 2026 02:40 |
| URI: | http://repository.its.ac.id/id/eprint/136344 |
Actions (login required)
![]() |
View Item |
