Rancang Bangun Website Pencarian Artikel Ilmiah Menggunakan Web Crawler Dan Reformulasi Query Berbasis Machine Learning

Virgiawan, Viera Tito (2026) Rancang Bangun Website Pencarian Artikel Ilmiah Menggunakan Web Crawler Dan Reformulasi Query Berbasis Machine Learning. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5026221096-Undergraduate_Thesis.pdf] Text
5026221096-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (6MB) | Request a copy

Abstract

Pencarian artikel ilmiah lintas repositori sering terkendala oleh fragmentasi sumber data dan ketidaksesuaian kosakata antara query pengguna dan istilah pada dokumen, sehingga hasil pencarian menjadi kurang relevan meskipun volume dokumen yang tersedia besar. Penelitian ini merancang dan membangun sebuah website pencarian artikel ilmiah bilingual (Indonesia dan Inggris) yang mengintegrasikan web crawler lintas repositori dengan modul reformulasi query berbasis machine learning. Sistem mengumpulkan metadata artikel dari tiga repositori terbuka, yaitu ArXiv, Semantic Scholar, dan DOAJ, kemudian menerapkan pipeline reformulasi query yang terdiri atas normalisasi, dekomposisi query berbasis dependency parsing, dan ekspansi Pseudo-Relevance Feedback. Setiap kandidat query yang dihasilkan ditelusuri melalui dua jalur retrieval yang berjalan berdampingan, yaitu pencarian sparse berbasis BM25 dan pencarian dense berbasis embedding multibahasa, sebelum digabungkan menggunakan Reciprocal Rank Fusion untuk menghasilkan satu peringkat akhir. Filosofi candidates-not-commitments diterapkan sebagai prinsip rancangan utama, di mana kesalahan pada satu kandidat dekomposisi tidak menghilangkan kandidat lain yang tetap valid, sehingga kesalahan tersebut diserap oleh mekanisme fusi peringkat tanpa menurunkan kualitas hasil akhir secara keseluruhan. Evaluasi dilakukan melalui pengujian unit, pengujian fungsional, evaluasi ablasi terhadap 47 query gold standard, dan evaluasi pencarian menggunakan metrik precision, recall, F1-score, NDCG, dan Mean Average Precision (MAP=0.451). Hasil ablasi menunjukkan bahwa dekomposisi query dan ekspansi PRF tidak menghasilkan perbedaan performa yang signifikan secara statistik dibandingkan baseline, konsisten dengan filosofi candidates-not-commitments yang memvalidasi ketahanan sistem terhadap kesalahan reformulasi individual. Performa retrieval yang lebih rendah pada domain tertentu, seperti pertanian dan perikanan, diatribusikan pada keterbatasan cakupan korpus, bukan pada kegagalan pipeline reformulasi atau retrieval.
====================================================================================================================================
Cross-repository scientific article search is often hindered by data source fragmentation and vocabulary mismatch between user queries and document terms, resulting in less relevant search results despite the large volume of available documents. This research designs and develops a bilingual (Indonesian and English) scientific article search website that integrates a cross-repository web crawler with a machine learning-based query reformulation module. The system collects article metadata from three open repositories, namely ArXiv, Semantic Scholar, and DOAJ, then applies a query reformulation pipeline consisting of normalization, dependency parsing-based query decomposition, and Pseudo-Relevance Feedback expansion. Each resulting query candidate is retrieved through two parallel retrieval paths, sparse retrieval based on BM25 and dense retrieval based on multilingual embeddings, before being merged using Reciprocal Rank Fusion to produce a single final ranking. The candidates-not-commitments philosophy is applied as the core design principle, in which an error in one decomposition candidate does not eliminate other candidates that remain valid, allowing such errors to be absorbed by the rank fusion mechanism without degrading overall result quality. Evaluation was conducted through unit testing, functional testing, an ablation study on 47 gold standard queries, and a search evaluation using precision, recall, F1-score, NDCG, and Mean Average Precision (MAP=0.451) metrics. Ablation results show that query decomposition and PRF expansion do not produce a statistically significant performance difference compared to the baseline, consistent with the candidates-not-commitments philosophy, which validates the system's resilience to individual reformulation errors. Lower retrieval performance in certain domains, such as agriculture and fisheries, is attributed to corpus coverage limitations rather than failure of the reformulation or retrieval pipeline.

Item Type: Thesis (Other)
Uncontrolled Keywords: BM25, Fusi Peringkat Timbal Balik, Machine Learning, Reformulasi Query, Web Crawler, BM25, Machine Learning, Query Reformulation, Reciprocal Rank Fusion, Web Crawler
Subjects: T Technology > T Technology (General)
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 57201-(S1) Undergraduate Thesis
Depositing User: Viera Tito Virgiawan
Date Deposited: 28 Aug 2026 00:30
Last Modified: 28 Aug 2026 00:30
URI: http://repository.its.ac.id/id/eprint/144393

Actions (login required)

View Item View Item