Analisis Strategi Retrieval (Chunking, Embedding, Dan Similarity Search) Dalam Arsitektur Retrieval Augmented Generation Untuk Pengolahan Dokumen Regulasi Energi Baru Terbarukan Di Indonesia

Maharani, Dewi (2026) Analisis Strategi Retrieval (Chunking, Embedding, Dan Similarity Search) Dalam Arsitektur Retrieval Augmented Generation Untuk Pengolahan Dokumen Regulasi Energi Baru Terbarukan Di Indonesia. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5026221046-Undergraduate_Thesis.pdf] Text
5026221046-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Dalam era transisi energi global, Indonesia menghadapi tantangan dalam akses informasi regulasi Energi Baru Terbarukan (EBT) akibat tumpang tindih kebijakan serta minimnya integrasi antar dokumen hukum. Kondisi ini mendorong pengembangan sistem tanya jawab berbasis Retrieval-Augmented Generation (RAG) dengan model generative Qwen2.5-7B-Instruct untuk meningkatkan efektivitas pencarian informasi regulasi. Penelitian ini mengoptimalkan pepiline retrieval melalui pengujian 12 kombinasi konfigurasi yang terdiri atas tiga metode chunking (fixed-length, overlap, dan semantic chunking), dua model embedding (intfloat/multilingual-e5-large dan firqaaa/indo-sentence-bert-base), serta dua strategi pemeringkatan hasil retrieval, yaitu dense retrieval dengan dan tanpa reranking menggunakan BAAI/bge-reranker-v2-m3. Dataset penelitin mencakup 100 dokumen regulasi sektor EBT (UU, PP, Perpres, Permen ESDM, dan regulasi terkait) periode 2005-2025 yang diproses menggunakan pdfplumber dan disimpan dalam vector database Milvus.Evaluasi dilakukan pada 50 pasangan query-expected answer menggunakan metrik cakupan token jawaban acuan, cakupan leksikal chunk, dan keselarasan semantic chunk, serta divalidasi melalui penilaian manual pada konfigurasi terbaik Evaluasi generasi jawaban menggunakan ROUGE-1, ROUGE-2, ROUGE-L, BERTScore, serta uji kesesuaian jawaban secara manual. Hasil penelitian menunjukkan bahwa konfigurasi Fixed-Length Chunking, Multilingual-e5-large, dan reranking (C1_E1_S2) memberikan performa terbaik secara konsisten, dengan persentase chunk relevan sebesar 64,40%, ROUGE-1 F1 sebesar 0,6031, BERTScore F1 sebesar 0,8149, serta 66% jawaban dinilai sepenuhnya sesuai. Temuan ini menujukkan bahwa pemilihan model embedding memiliki pengaruh lebih dominan dibandingkan metode chunking, sementara reranking meningkatkan relevansi kontekstual dan ketepatan jawaban akhir. Penelitian ini berkontribusi pada pengembangan metodologi evaluasi RAG pada dokumen hukum berbahasa Indonesia tanpa ketergantungan pada qrels, serta memberikan manfaat praktis dalam mendukung pengambilan Keputusan pada sektor energi terbarukan.
=====================================================================================================================================
In the era of global energy transition, Indonesia faces challenges in accessing regulatory information on New and Renewable Energy (NRE) due to overlapping policies and minimal integration across legal documents. This condition motivated the development of a question-answering system based on Retrieval-Augmented Generation (RAG) using the generative model Qwen2.5-7B-Instruct to improve the effectiveness of regulatory information retrieval. This research optimizes the retrieval pipeline by testing 12 configuration combinations consisting of three chunking methods (fixed-length, overlap, and semantic chunking), two embedding models (intfloat/multilingual-e5-large and firqaaa/indo-sentence-bert-base), and two retrieval re-ranking strategies, namely dense retrieval with and without reranking using BAAI/bge-reranker-v2-m3. The research dataset comprises 100 NRE sector regulatory documents (laws, government regulations, presidential regulations, ministerial regulations of the Ministry of Energy and Mineral Resources, and related regulations) from the period 2005–2025, processed using pdfplumber and stored in the Milvus vector database. Evaluation was conducted on 50 query–expected answer pairs using the expected answer token coverage, chunk lexical coverage, and chunk semantic alignment metrics, and validated through manual assessment of the best-performing configuration. Answer generation evaluation employed ROUGE-1, ROUGE-2, ROUGE-L, BERTScore, and manual answer-correctness assessment. The results show that the configuration combining Fixed-Length Chunking, Multilingual-E5-Large, and reranking (C1_E1_S2) consistently delivered the best performance, achieving a relevant chunk percentage of 64.40%, a ROUGE-1 F1 score of 0.6031, a BERTScore F1 of 0.8149, and 66% of answers rated as fully appropriate. These findings indicate that the choice of embedding model has a more dominant influence than the chunking method, while reranking improves contextual relevance and the accuracy of the final answer. This research contributes to the development of RAG evaluation methodology for Indonesian-language legal documents without reliance on qrels, and provides practical benefits in supporting decision-making in the renewable energy sector.

Item Type: Thesis (Other)
Uncontrolled Keywords: Chunking, Embedding, Similarity Search, RAG, EBT, Chunking, Embedding, Similarity Search, RAG, EBT.
Subjects: T Technology > T Technology (General) > T57.5 Data Processing
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 57201-(S1) Undergraduate Thesis
Depositing User: Dewi Maharani
Date Deposited: 31 Jul 2026 01:47
Last Modified: 31 Jul 2026 01:47
URI: http://repository.its.ac.id/id/eprint/139515

Actions (login required)

View Item View Item