Optimasi LLM pada Prompt-Engineering dalam Penyelesaian Soal Matematika dengan Multi-Stage dan Knowledge Integration untuk Pembelajaran Siswa

Rokhim, Imam Fadhkur (2026) Optimasi LLM pada Prompt-Engineering dalam Penyelesaian Soal Matematika dengan Multi-Stage dan Knowledge Integration untuk Pembelajaran Siswa. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 6025241010-Master_Thesis.pdf] Text
6025241010-Master_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (5MB) | Request a copy

Abstract

Kemampuan penalaran matematis pada Large Language Models (LLM) masih menghadapi tantangan berupa kesalahan logika, halusinasi, dan keterbatasan pengetahuan dalam penyelesaian soal, sehingga berpotensi menurunkan keandalan sebagai sarana pendukung belajar. Penelitian ini mengusulkan framework Multi-Stage Knowledge Integration (MSKI) untuk meningkatkan kualitas penyelesaian soal matematika serta menghasilkan penjelasan yang lebih mudah dipahami siswa. Framework ini mengintegrasikan Systematic Decomposition, Contextual Knowledge Retrieval, Knowledge-Infused Inference, Feedback Driven Refinement, Fuzzy Confidence Aggregation (FCA), dan Final Pedagogical Refinement dalam satu alur terpadu untuk memecah masalah kompleks, memanfaatkan materi pendukung yang sesuai, melakukan perbaikan bertahap, dan memilih solusi paling andal melalui agregasi berbasis fuzzy.
Pengujian dilakukan menggunakan empat model LLM, yaitu Gemini 2.5 Flash, DeepSeek-Chat, GPT-4o Mini, dan Qwen2.5-7B-Instruct. Evaluasi memanfaatkan dua dataset, yakni Indonesian Mathematical Books yang berisi 4.925 soal berbahasa Indonesia dari jenjang SD hingga SMA serta MathOdyssey yang mencakup 389 soal tingkat SMA dan universitas. Kinerja dibandingkan dengan Chain of Thought (CoT), Decompose, Analyze, and Rethink (DeAR), serta Progressive Back-Translation (PBT) menggunakan metrik Accuracy, BERTScore, panjang penalaran, konsumsi token, waktu inferensi, dan penilaian manusia.
Hasil eksperimen menunjukkan bahwa varian MSKI-FCA memperoleh performa terbaik pada aspek akurasi. Pada Indonesian Mathematical Books, metode tersebut mencapai nilai 90,39% menggunakan Gemini 2.5 Flash, sedangkan pada MathOdyssey memperoleh 81,61%. Penilaian guru dan siswa juga memperlihatkan kualitas penjelasan yang baik dari sisi validitas logika, kemudahan dipahami, serta kegunaan untuk belajar mandiri. Dengan demikian, MSKI berpotensi menjadi pendekatan efektif untuk mendukung pembelajaran matematika berbasis kecerdasan buatan.
=====================================================================================================================================
Mathematical reasoning in Large Language Models (LLMs) still facechallenges such as logical errors, hallucinations, and limited knowledge in problemsolving, potentially reducing their reliability as a learning tool. This researchproposes a Multi-Stage Knowledge Integration (MSKI) framework to improve thequality of math problem solving and produce explanations that are more easilyunderstood by students. This framework integrates Systematic Decomposition,Contextual Knowledge Retrieval, Knowledge-Infused Inference, Feedback-DrivenRefinement, Fuzzy Confidence Aggregation (FCA), and Final PedagogicalRefinement into a single, integrated flow to break down complex problems, utilizeappropriate supporting materials, make incremental improvements, and select themost reliable solution through fuzzy-based aggregation.Testing was conducted using four LLM models: Gemini 2.5 Flash,DeepSeek-Chat, GPT-4o Mini, and Qwen2.5-7B-Instruct. The evaluation utilizedtwo datasets: the Indonesian Mathematical Books, which contains 4,925Indonesian-language problems from elementary to high school levels, andMathOdyssey, which includes 389 problems from high school and university levels.Performance was compared with Chain of Thought (CoT), Decompose, Analyze,and Rethink (DeAR), and Progressive Back-Translation (PBT) using the metricsAccuracy, BERTScore, reasoning length, token consumption, inference time, andhuman judgment.The experimental results showed that the MSKI-FCA variant performedbest in terms of accuracy. On the Indonesian Mathematical Books, the methodachieved a score of 90.39% using Gemini 2.5 Flash, while on MathOdyssey itachieved 81.61%. Teacher and student assessments also demonstrated goodexplanation quality in terms of logical validity, ease of understanding, and usabilityfor independent learning. Thus, MSKI has the potential to be an effective approachto supporting AI-based mathematics learning.

Item Type: Thesis (Masters)
Uncontrolled Keywords: Fuzzy Confidence Aggregation, Knowledge Integration, Large Language Models, Multi-Stage Reasoning, Prompt Engineering, Penalaran Matematika, Progressive Back-Translation, Keywords: Fuzzy Confidence Aggregation, Knowledge Integration, LargeLanguage Models, Mathematical Reasoning, Multi-Stage Reasoning, PromptEngineering, Progressive Back-Translation
Subjects: L Education > LB Theory and practice of education > LB1044.87 Internet in education (e-learning). Virtual reality in education.
T Technology > T Technology (General)
T Technology > T Technology (General) > T57.5 Data Processing
T Technology > T Technology (General) > T57.6 Operations research--Mathematics. Goal programming
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: Imam Fadhkur Rokhim
Date Deposited: 28 Jul 2026 03:20
Last Modified: 28 Jul 2026 03:20
URI: http://repository.its.ac.id/id/eprint/138333

Actions (login required)

View Item View Item