Ramadhani, Reynaldi Neo (2026) Pengembangan Sistem Ekstraksi Environmental, Social, and Governance (ESG) Berbasis Retrieval Augmented Generation pada Dokumen Multi-Modal. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221265-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (8MB) | Request a copy |
Abstract
Informasi Environmental, Social, and Governance (ESG) penting untuk menilai keberlanjutan perusahaan, namun datanya tersebar pada berbagai bagian dokumen dan disajikan dalam bentuk teks, tabel, maupun gambar, sehingga sulit diekstraksi secara sistematis. Penelitian ini bertujuan merancang dan mengembangkan sistem ekstraksi data ESG dari dokumen perusahaan multi-modal menggunakan pendekatan Retrieval Augmented Generation (RAG) untuk pencarian evidence dan Vision-Language Model (VLM) untuk merepresentasikan informasi visual, serta mengevaluasi kinerjanya terhadap ground truth. Sistem dibangun melalui pipeline ingestion, retrieval, dan ekstraksi. Pada tahap ingestion, dokumen PDF dipecah menjadi chunk teks dan elemen visual diubah menjadi caption menggunakan VLM, lalu direpresentasikan sebagai embedding. Tahap retrieval menggunakan query builder berbasis framework indikator, vector search, Reciprocal Rank Fusion (RRF), dan heuristic reranker, sedangkan ekstraksi memanfaatkan Large Language Model (LLM) dengan prompt berbasis evidence. Evaluasi dilakukan melalui studi kasus terhadap 3 dokumen perusahaan (150 halaman), 81 indikator ESG, dan 97 baris ground truth menggunakan model keluarga Qwen. Konfigurasi terbaik memperoleh Hit@5 0,7160, MRR 0,6650, akurasi penentuan disclosure (AccDC) 0,7407, dan menghasilkan output untuk seluruh indikator (Prediction Coverage 1,0000). Sistem mampu menjalankan ekstraksi ESG secara end-to-end dengan performa tertinggi pada evidence teks dan terendah pada evidence gambar. Tantangan utama yang masih dihadapi adalah retrieval evidence visual dan ekstraksi nilai numerik, yang ditunjukkan oleh rendahnya akurasi ekstraksi nilai (AccDE 0,4444).
====================================================================================================================================
Environmental, Social, and Governance (ESG) information is important for assessing corporate sustainability, yet such data is scattered across various parts of company documents and presented as text, tables, and images, making its systematic extraction difficult. This research aims to design and develop an ESG data extraction system from multi-modal company documents using the Retrieval Augmented Generation (RAG) approach for evidence retrieval and a Vision-Language Model (VLM) to represent visual information, and to evaluate its performance against the ground truth. The system is built through an ingestion, retrieval, and extraction pipeline. In the ingestion stage, PDF documents are split into text chunks and visual elements are converted into captions using a VLM, then represented as embeddings. The retrieval stage employs an indicator-framework-based query builder, vector search, Reciprocal Rank Fusion (RRF), and a heuristic reranker, while the extraction stage uses a Large Language Model (LLM) with evidence-based prompts. The evaluation was conducted through a case study on 3 company documents (150 pages), 81 ESG indicators, and 97 ground truth rows using the Qwen model family. The best configuration achieved a Hit@5 of 0.7160, an MRR of 0.6650, a disclosure determination accuracy (AccDC) of 0.7407, and produced output for all indicators (Prediction Coverage of 1.0000). The system is able to perform ESG extraction end-to-end with the highest performance on text-based evidence and the lowest on image-based evidence. The main challenges remain in retrieving visual evidence and extracting numerical values, as reflected by the relatively low data extraction accuracy (AccDE 0.4444).
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Dokumen Multi-Modal,Ekstraksi ESG,Large Language Model,Retrieval Augmented Generation,Vision-Language Model, ESG Extraction,Large Language Model,Multi-Modal Document,Retrieval Augmented Generation,Vision-Language Model |
| Subjects: | Q Science > QA Mathematics > QA336 Artificial Intelligence Q Science > QA Mathematics > QA76 Computer software Q Science > QA Mathematics > QA76.758 Software engineering Q Science > QA Mathematics > QA76.9.D343 Data mining. Querying (Computer science) |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Reynaldi Neo Ramadhani |
| Date Deposited: | 28 Jul 2026 06:37 |
| Last Modified: | 28 Jul 2026 06:37 |
| URI: | http://repository.its.ac.id/id/eprint/138634 |
Actions (login required)
![]() |
View Item |
