Sudana, Pranaditya Tri Jyotista Vavitram Putra (2026) Pengaruh Metode Text Chunking Terhadap Kualitas Konteks dan Performa Sistem RAG Untuk Help Desk. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5024221052-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (8MB) |
Abstract
Perkembangan chatbot berbasis kecerdasan buatan sebagai asisten help desk membutuhk-
an arsitektur Retrieval-Augmented Generation (RAG) untuk memberikan respons yang akurat
berdasarkan dokumen institusi. Namun, kinerja sistem RAG sering kali terhambat oleh metode
segmentasi teks konvensional (fixed-size chunking) yang memotong informasi secara statis.
Pendekatan arbitrer ini berisiko menghilangkan konteks semantik dan menurunkan kualitas
dokumen yang diambil (retrieval). Penelitian ini bertujuan untuk melakukan studi komparatif
terhadap enam metode text chunking, yaitu Fixed-Size, Semantic, Propositional, Recursive,
Topic-Aware, dan Multi-Granular Chunking, guna mengevaluasi pengaruhnya terhadap kualitas
konteks dan performa sistem RAG pada domain dokumen administrasi Teknik Komputer ITS.
Pengujian dilakukan menggunakan metode Research and Development (RD) dengan mengevalu-
asi efektivitas retrieval menggunakan metrik NDCG dan Recall, serta mengukur kualitas jawaban
dari Large Language Model (LLM) melalui framework evaluasi DeepEval (Context Precision,
Context Relevancy, Faithfulness, dan Answer Relevancy). Hasil pendahuluan menunjukkan
adanya trade-off (kompromi) antara efektivitas penemuan dokumen dan latensi komputasi sistem.
Pemotongan berbasis semantik dan proposisi terbukti memberikan keseimbangan yang lebih baik
dalam mempertahankan koherensi makna dibandingkan pemotongan statis. Hasil dari penelitian
ini diharapkan dapat menjadi acuan arsitektur dalam menentukan metode segmentasi yang paling
optimal untuk menunjang performa help desk assistant.
=====================================================================
The development of AI-based chatbots as help desk assistants requires a Retrieval-Augmented
Gen- eration (RAG) architecture to provide accurate responses based on institutional documents.
However, the performance of RAG systems is often hindered by conventional text segmentation
methods (fixed-size chunking) that cut information statically. This arbitrary approach risks
losing semantic context and degrading the quality of retrieved documents. This research aims
to conduct a comparative study of six text chunking methods: Fixed-Size, Semantic, Proposi-
tional, Recursive, Topic-Aware, and Multi- Granular Chunking, to evaluate their impact on
context quality and RAG system performance within the domain of administrative documents
at Computer Engineering ITS. The evaluation is conducted using a Research and Development
(RD) approach, measuring retrieval effectiveness via NDCG and Recall metrics, and assessing
the Large Language Model (LLM) answer quality through the DeepEval frame- work (Context
Precision,Context Relevancy, Faithfulness, and Answer Relevancy). Preliminary results indicate
a trade-off between document retrieval effectiveness and computational latency. Semantic and
propositional-based chunking demonstrate a better balance in preserving meaning coherence
compared to static chunking. The results of this research are expected to serve as an architectural
reference in determining the most optimal segmentation method to support the performance of a
help desk assistant
Actions (login required)
![]() |
View Item |
