Named-Entity Recognition untuk Pengenalan Karakter dan Penokohan Cerita Berbahasa Sunda dengan Large Language Model

Amrullah, Muhammad Syiarul (2026) Named-Entity Recognition untuk Pengenalan Karakter dan Penokohan Cerita Berbahasa Sunda dengan Large Language Model. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 6025232012-Master_Thesis.pdf] Text
6025232012-Master_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Identifikasi tokoh beserta perannya dalam teks berbahasa Sunda masih menghadapi tantangan. Bahasa Sunda menunjukkan variasi leksikal, fonologis, dan morfologis yang cukup besar di berbagai daerah. Variasi ini mempersulit analisis secara manual karena membutuhkan sumber daya besar, dan menurunkan efektivitas metode ekstraksi informasi secara konvensional. Named Entity Recognition berbasis Large Language Models dikembangkan untuk secara mengidentifikasi tokoh beserta perannya dalam teks naratif Sunda. Selain itu, penelitian ini juga mengevaluasi pengaruh integrasi peran label semantik terhadap performa model. Metode yang diusulkan menggunakan LLMs berbasis transformer untuk menghasilkan embedding kontekstual dari tokenisasi subword. Embedding ini kemudian digabungkan untuk memperkaya konteks semantik saat klasifikasi. Korpus yang digunakan terdiri atas 98 teks naratif Sunda, termasuk fabel, legenda, dan dongeng rakyat. Proses anotasi label tokoh dan peran dibagi menjadi dua yaitu 68 teks secara manual, 30 teks dianotasi dengan pseudo-label. Hasil eksperimen menunjukkan bahwa Large Language Models dengan peran label semantik mencapai skor F1 sebesar 70,45%. Performa ini melampaui model tradisional seperti CRF dan HMM, masing-masing sekitar 5%, dan peningkatan performa dari baseline dengan menyediakan hubungan semantik yang diperkaya dalam kalimat. Pengujian hipotesis statistik Uji paired bootstrap dan Uji McNemar's) mengonfirmasi bahwa peningkatan ini signifikan secara statistik (p < 0,001). Temuan ini menunjukkan bahwa LLM yang terintegrasi dengan peran label semantik memberikan pendekatan yang lebih efektif untuk identifikasi tokoh beserta perannya dalam teks naratif berbahasa Sunda. Hasilnya, memperlihatkan bahwa pendekatan ini mengungguli metode statistik konvensional dan baseline. Selain itu, untuk meningkatkan penelitian ini diperlukan integrasi coreference resolution guna memperkaya informasi pada entitas OBJ yang bersifat umum. Lalu, dibutuhkan penelitian lebih lanjut terhadap Semantic Role Labelling secara otomatis untuk integrasi dengan mulus.
=======================================================================================================================================
Identifying characters and their roles in Sundanese texts remains a challenging task due to the substantial lexical, phonological, and morphological variations that exist across different regions. These variations complicate manual analysis by requiring extensive resources and reduce the effectiveness of conventional information extraction methods. To address this issue, this study develops a Named Entity Recognition (NER) approach based on Large Language Models (LLMs) to automatically identify characters and their associated roles in Sundanese narrative texts. In addition, the study investigates the impact of integrating semantic role labels on model performance. The proposed method employs transformer-based LLMs to generate contextual embeddings from subword tokenization. These embeddings are combined with semantic role representations through embedding concatenation to enrich semantic context during the classification process. The corpus consists of 98 Sundanese narrative texts, including fables, legends, and folk tales. The annotation process was conducted in two stages: 68 texts were manually annotated, while 30 texts were annotated using a pseudo-labeling approach.
Experimental results demonstrate that the LLM integrated with semantic role labels achieves an F1-score of 70.45%. This performance surpasses traditional models such as Conditional Random Fields (CRF) and Hidden Markov Models (HMM) by approximately 5%, respectively. Statistical hypothesis testing (paired bootstrap test and McNemar's test) confirms that these improvements are statistically significant (p < 0.001). The improvement is attributed to the enhanced semantic relationships provided by semantic role information within sentences. These findings indicate that integrating semantic role labels into LLMs offers a more effective approach for identifying characters and their roles in Sundanese narrative texts, outperforming both conventional statistical methods and baseline models. Future research should incorporate coreference resolution to enrich information related to generic object entities and explore automatic Semantic Role Labeling techniques to enable more seamless integration within the proposed framework.

Item Type: Thesis (Masters)
Uncontrolled Keywords: Model Bahasa Besar, Pengenalan Entitas Bernama, Pengenalan Karakter, Pengenalan Peran Karakter, Peran Label Semantik, Character Identification, Large Language Model, Named Entity Recognition, Role Identification, Semantic Role Label
Subjects: T Technology > T Technology (General) > T57.5 Data Processing
T Technology > T Technology (General) > T57.8 Nonlinear programming. Support vector machine. Wavelets. Hidden Markov models.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: Muhammad Syiarul Amrullah
Date Deposited: 29 Jul 2026 04:05
Last Modified: 29 Jul 2026 04:05
URI: http://repository.its.ac.id/id/eprint/139505

Actions (login required)

View Item View Item