Arsitektur Hybrid Retrieval Dengan Generative Query Expansion Berbasis LLM Untuk Sistem Penemuan Pakar Berbahasa Indonesia

Zam Zami, Nuril Qolbi (2026) Arsitektur Hybrid Retrieval Dengan Generative Query Expansion Berbasis LLM Untuk Sistem Penemuan Pakar Berbahasa Indonesia. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221296-Undergraduate_Thesis.pdf] Text
5025221296-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Sistem penemuan pakar (expert finding) di lingkungan akademik menghadapi dua tantangan utama, yaitu ketidakcocokan kosakata (vocabulary mismatch) antara kueri pengguna dan profil pakar, serta keterbatasan model semantik dalam pencocokan istilah secara presisi (exact match). Di Institut Teknologi Sepuluh Nopember (ITS), tantangan tersebut ditemui pada data 863 profil dosen yang keahliannya tersebar di 42 program studi. Penelitian ini mengusulkan arsitektur hybrid retrieval yang menggabungkan pencarian leksikal BM25 dengan pencarian semantik IndoBERT, diperkuat dengan Generative Query Expansion (GQE) berbasis dua large language model open-source (Qwen2-7B-Instruct dan Meta-LLaMA-3-8B-Instruct), dengan seluruh parameter sistem ditentukan melalui hyperparameter tuning bertahap. Evaluasi dilakukan pada korpus 863 profil dosen ITS dan 148 kueri uji bertipe single-relevant, menggunakan metrik Mean Reciprocal Rank (MRR), Hits@10, Hits@20, dan Normalized Discounted Cumulative Gain (nDCG). Pada konfigurasi parameter hasil tuning, model Hybrid+GQE-Qwen2 memimpin performa secara konsisten di seluruh metrik evaluasi, dengan MRR sebesar 0,1169 (meningkat 35,61% dari baseline BM25), Hits@10 sebesar 25,00% (meningkat 32,14%), Hits@20 sebesar 30,41% (meningkat 32,39%), serta nDCG@10 sebesar 0,1455 dan nDCG@20 sebesar 0,1595, sementara model Hybrid+GQE-LLaMA3 menunjukkan performa yang berdekatan (MRR 0,1148, peningkatan relatif 33,18%). Peningkatan tersebut konsisten teramati pada kedua model GQE dibandingkan BM25, IndoBERT, dan Hybrid tanpa GQE, menunjukkan kontribusi integrasi Generative Query Expansion dalam arsitektur hybrid retrieval terhadap akurasi temu kembali pakar pada korpus berbahasa Indonesia.
==============================================================================================================================
Expert finding systems in academic settings face two central challenges: vocabulary mismatch between user queries and expert profiles, and the limited ability of semantic models to perform precise exact-term matching. At Institut Teknologi Sepuluh Nopember (ITS), these challenges appear in a dataset of 863 lecturer profiles whose expertise spans 42 study programs. This study proposes a hybrid retrieval architecture that combines lexical retrieval (BM25) with semantic retrieval (IndoBERT), enhanced by Generative Query Expansion (GQE) based on two open-source large language models (Qwen2-7B-Instruct and Meta-LLaMA-3-8B-Instruct), with all system parameters determined through staged hyperparameter tuning. Evaluation was conducted on a corpus of 863 ITS lecturer profiles and 148 single-relevant test queries, using Mean Reciprocal Rank (MRR), Hits@10, Hits@20, and Normalized Discounted Cumulative Gain (nDCG) as evaluation metrics. Under the tuned parameter configuration, Hybrid+GQE-Qwen2 led performance consistently across all evaluation metrics, achieving an MRR of 0.1169 (a 35.61% increase over the BM25 baseline), Hits@10 of 25.00% (a 32.14% increase), Hits@20 of 30.41% (a 32.39% increase), along with an nDCG@10 of 0.1455 and an nDCG@20 of 0.1595, while Hybrid+GQE-LLaMA3 showed closely comparable performance (MRR 0.1148, a 33.18% relative increase). These improvements were observed consistently across both GQE-based models relative to BM25, IndoBERT, and the hybrid model without GQE, indicating the contribution of integrating Generative Query Expansion within the hybrid retrieval architecture to expert retrieval accuracy on an Indonesian-language corpus.

Item Type: Thesis (Other)
Uncontrolled Keywords: Expert Finding, Generative Query Expansion, Hybrid Retrieval, IndoBERT, Large Language Models ============================================================================================================================== Expert Finding, Generative Query Expansion, Hybrid Retrieval, IndoBERT, Large Language Models
Subjects: Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
Q Science > QA Mathematics > QA76.9.D343 Data mining. Querying (Computer science)
Q Science > QA Mathematics > QA76.9.I58 Recommender systems (Information filtering)
Z Bibliography. Library Science. Information Resources > Z665 Library Science. Information Science
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Nuril Qolbi Zam Zami
Date Deposited: 28 Jul 2026 01:26
Last Modified: 28 Jul 2026 01:26
URI: http://repository.its.ac.id/id/eprint/138159

Actions (login required)

View Item View Item