Asmani, Anani (2026) Integrasi Retrieval-Augmented Generation Dan Reinforcement Learning Untuk Meningkatkan Performa Mathematical Question Answering Berbasis Large Language Model. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6025241008-Master_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (1MB) | Request a copy |
Abstract
Evaluasi pembelajaran matematika yang berfokus pada pertanyaan kompleks sangat penting untuk mengukur kemampuan berpikir kritis, namun proses penilaian jawaban esai ini membutuhkan ketelitian tinggi dan waktu yang lama bagi guru. Large Language Model (LLM) memiliki potensi untuk membantu mengotomatisasi proses penilaian tersebut. Penelitian ini mengusulkan pendekatan baru dengan menggabungkan RAG dan Reinforcement Learning (RL) dalam kerangka kerja Invoke-Verify-Inject. RAG berfungsi sebagai Knowledge Invoker untuk membatasi ruang pencarian jawaban agar sesuai dengan konteks. Sedangkan RL berfungsi mengevaluasi setiap langkah penyelesaian dan memberi reward pada setiap langkah penalaran yang valid secara matematis. Sistem dirancang bekerja secara iteratif (looping), memungkinkan langkah yang tidak logis diperbaiki secara otomatis sebelum menghasilkan jawaban akhir. Eksperimen dilakukan menggunakan 4.925 soal matematika dari jenjang SD, SMP, dan SMA. Empat metode dibandingkan, yaitu model dasar, LLM dengan RAG, LLM dengan RL, dan gabungan RAG-RL. Hasil penelitian menunjukkan metode usulan unggul pada seluruh metrik. Pada Factual Accuracy Assessment (FAA), akurasi meningkat dari 0,089–0,116 pada baseline menjadi 0,6434–0,6778 pada metode usulan, dengan DeepSeek-Chat R1 sebagai model terbaik (FAA 0,6778). Metode usulan juga mencatat nilai tertinggi pada BERTScore (0,8163–0,8351) dan Consistency (0,9152–0,9329). Uji paired t-test menunjukkan keunggulan metode usulan signifikan secara statistik (p < 0,001) pada seluruh perbandingan. Pada evaluasi kualitatif oleh lima evaluator (GPT, DeepSeek, Grok) terhadap aspek relevansi, answerability, dan reliabilitas, DeepSeek-Chat R1 unggul pada soal HOTS (relevansi 80%, answerability 84%, reliabilitas 82%), sedangkan GPT-4o-Mini unggul pada soal LOTS dengan answerability mencapai 88% sebagai skor tertinggi di seluruh hasil evaluasi.
==================================================================================================================================
Mathematics learning evaluations that focus on complex questions are crucial for measuring critical thinking skills, but the process of assessing essay answers requires high precision and is time-consuming for teachers. Large Language Models (LLMs) have the potential to help automate this assessment process. This study proposes a new approach by combining RAG and Reinforcement Learning (RL) within the Invoke-Verify-Inject framework. RAG functions as a Knowledge Invoker to limit the search space for answers to suit the context. RL, on the other hand, evaluates each step of the solution and rewards each mathematically valid reasoning step. The system is designed to work iteratively (looping), allowing illogical steps to be automatically corrected before producing the final answer. Experiments were conducted using 4,925 mathematics problems from elementary, middle, and high school levels. Four methods were compared: the baseline model, LLM with RAG, LLM with RL, and a combination of RAG and RL. The results showed that the proposed method excelled across all metrics. In the Factual Accuracy Assessment (FAA), accuracy increased from 0.089–0.116 in the baseline to 0.6434–0.6778 in the proposed method, with DeepSeek-Chat R1 as the best model (FAA 0.6778). The proposed method also recorded the highest scores in BERTScore (0.8163–0.8351) and Consistency (0.9152–0.9329). The paired t-test showed the superiority of the proposed method was statistically significant (p < 0.001) in all comparisons. In the qualitative evaluation by five evaluators (GPT, DeepSeek, Grok) on the aspects of relevance, answerability, and reliability, DeepSeek-Chat R1 excelled on HOTS questions (80% relevance, 84% answerability, 82% reliability), while GPT-4o-Mini excelled on LOTS questions with 88% answerability as the highest score in all evaluation results.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | Large Language Model (LLM), Penalaran Matematis, Reinforcement Learning (RL), Retrieval-Augmented Generation (RAG), Verify, Large Language Model (LLM), Mathematical Reasoning, Reinforcement Learning (RL), Retrieval-Augmented Generation (RAG), Verify |
| Subjects: | Q Science > QA Mathematics > QA336 Artificial Intelligence Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science) |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis |
| Depositing User: | Anani Asmani |
| Date Deposited: | 30 Jul 2026 00:34 |
| Last Modified: | 30 Jul 2026 00:34 |
| URI: | http://repository.its.ac.id/id/eprint/139688 |
Actions (login required)
![]() |
View Item |
