Pengembangan Metode Ekstraksi Informasi dan Pemberian Rekomendasi pada Laporan Psikologis Millon Clinical Multiaxial Inventory IV (MCMI-IV) menggunakan Named Entity Recognition (NER) dan Retrieval Augmented Generation (RAG)

Kemaputra, Anas Ghifari (2026) Pengembangan Metode Ekstraksi Informasi dan Pemberian Rekomendasi pada Laporan Psikologis Millon Clinical Multiaxial Inventory IV (MCMI-IV) menggunakan Named Entity Recognition (NER) dan Retrieval Augmented Generation (RAG). Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5026221155-Undergraduate_Thesis.pdf] Text
5026221155-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (5MB) | Request a copy

Abstract

Meningkatnya prevalensi gangguan mental di Indonesia menuntut layanan psikologis yang komunikatif bagi klien. Millon Clinical Multiaxial Inventory-IV (MCMI-IV) merupakan alat tes kesehatan mental yang menghasilkan skor sangat teknis dan sulit diinterpretasikan. Para psikolog mulai memanfaatkan AI (Artificial Intelligence) untuk menginterpretasikan skor tersebut, namun penggunaan AI tanpa arsitektur khusus kerap menghasilkan interpretasi yang tidak terstandarisasi, rentan halusinasi data, dan bersifat kaku. Penelitian ini mengatasi permasalahan tersebut melalui pengembangan model NLP yang mengintegrasikan Named Entity Recognition (NER) dan Retrieval-Augmented Generation (RAG). Model NER mengekstraksi entitas klinis berupa gejala, kekuatan, dan tantangan dari data mentah, kemudian merangkumnya menjadi narasi terstruktur, sementara RAG memperkaya bagian rekomendasi menjadi narasi yang hangat dan mudah dicerna klien. Penelitian menggunakan tiga puluh laporan klinis nyata sebagai basis pengetahuan, membandingkan kinerja NER berbasis BERT, IndoBERT, dan BioBERT terhadap baseline LLM Gemini menggunakan F1-Score, serta mengukur kualitas narasi melalui ROUGE, BLEU, METEOR, dan evaluasi LLM as a Judge. Hasil menunjukkan IndoBERT sebagai model NER paling optimal dengan F1-Score Partial Match 60,61%, mengungguli Gemini (39,75%), BERT (37,45%), dan BioBERT (27,86%). RAG meningkatkan Recall secara signifikan dengan ROUGE-1 Recall sebesar 0,6538 dibandingkan Baseline yang hanya 0,4500, serta peningkatan METEOR menjadi 0,3042, mengindikasikan ringkasan yang lebih komprehensif dan kaya secara semantik. Meskipun evaluasi LLM as a Judge menunjukkan Baseline sedikit unggul pada Faithfulness (9,64 berbanding 8,07) dan Answer Relevance (9,86 berbanding 9,43), validasi expert (psikolog) menyimpulkan keluaran RAG lebih unggul karena lebih eksplanatif, menyeluruh, dan valid secara klinis. Inovasi ini diharapkan meningkatkan literasi kesehatan mental klien melalui penyampaian hasil yang lebih manusiawi dan akurat.
======================================================================================================================================
The rising prevalence of mental disorders in Indonesia demands psychological services that are communicative with clients. The Millon Clinical Multiaxial Inventory-IV (MCMI-IV) is a mental health assessment tool that produces highly technical scores that are difficult to interpret. Psychologists have begun using AI to interpret these scores; however, the use of AI without a specialized architecture often results in non-standardized interpretations that are prone to data hallucinations and lack flexibility. This study addresses these issues by developing an NLP model that integrates Named Entity Recognition (NER) and Retrieval-Augmented Generation (RAG). The NER model extracts clinical entities such as symptoms, strengths, and challenges from raw data and summarizes them into a structured narrative, while RAG enriches the recommendation section with warm, client-friendly narratives. The study used thirty real clinical reports as a knowledge base, comparing the performance of BERT, IndoBERT, and BioBERT-based NER models against the Gemini LLM baseline using the F1-Score, and measuring narrative quality through ROUGE, BLEU, METEOR, and LLM as a Judge evaluation. The results show that IndoBERT is the most optimal NER model with a weighted Partial Match F1-Score of 60.61%, outperforming Gemini (39.75%), BERT (37.45%), and BioBERT (27.86%). RAG significantly improves Recall, with a ROUGE-1 Recall of 0.6538 compared to the Baseline of 0.4500, and an increase in METEOR to 0.3042, indicating a more comprehensive and semantically richer summary. Although the LLM as a Judge evaluation showed the Baseline to be slightly superior in Faithfulness (9.64 against 8.07) and Answer Relevance (9.86 against 9.43), expert validation (by psychologists) concluded that RAG output was superior because it was more explanatory, comprehensive, and clinically valid. This innovation is expected to improve clients mental health literacy through the delivery of more human-centered and accurate results.

Item Type: Thesis (Other)
Uncontrolled Keywords: MCMI-IV, NER, RAG, NLP, Psikologi, MCMI-IV, NER, RAG, NLP, Psychology
Subjects: Q Science > QA Mathematics > QA336 Artificial Intelligence
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 57201-(S1) Undergraduate Thesis
Depositing User: Anas Ghifari Kemaputra
Date Deposited: 24 Jul 2026 01:16
Last Modified: 24 Jul 2026 01:16
URI: http://repository.its.ac.id/id/eprint/136835

Actions (login required)

View Item View Item