Ekstraksi Entitas Kriminalitas Menggunakan Named Entity Recognition berbasis Transformer Pada Pengelompokan Berita Daring

Wijaya, I Putu Raditya Partha (2026) Ekstraksi Entitas Kriminalitas Menggunakan Named Entity Recognition berbasis Transformer Pada Pengelompokan Berita Daring. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221210-Undergraduate_Thesis.pdf] Text
5025221210-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Kenaikan tren kriminalitas di Indonesia dan keterbatasan akses data resmi tingkat insiden mendorong kebutuhan sistem analitik yang mampu mengekstraksi informasi peristiwa kriminal secara otomatis dari berita daring. Penelitian ini mengembangkan pipeline awal sampai akhir yang mengubah narasi berita kriminal berbahasa Indonesia menjadi pemetaan kerentanan terstruktur yang memetakan cluster kejahatan terhadap tipe lokasi fisik, diperkaya dengan konteks tipe pelaku dan objek sasaran. Korpus sebanyak 13.937 artikel dikumpulkan dari detik.com pada periode 1 Januari 2023 hingga 31 Desember 2024 dan dibersihkan melalui normalisasi Unicode serta penghapusan boilerplate. Sebanyak 400 artikel diambil melalui stratified sampling dan dianotasi secara manual menggunakan skema BIO empat entitas: CRIME_ACT, LOC_TYPE, PERPETRATOR_TYPE, dan TARGET_ITEM pada tingkat karakter. Pelatihan NER dilakukan dalam dua tahap: Tahap 1 menentukan model BERT Bahasa Indonesia yang paling optimal dari empat kandidat (IndoBERT-IndoLEM, IndoBERT-IndoNLU, CahyaBERT, NusaBERT), dengan IndoBERT-IndoLEM memberikan performa terbaik dengan F1 val sebesar 0.6830; Tahap 2 menerapkan Bayesian hyperparameter optimization via Optuna yang meningkatkan F1 val menjadi 0,7353 dengan F1 test akhir sebesar 0.6658. Inferensi dua model dilakukan di seluruh korpus dengan model fine-tuned untuk deteksi span dan model base pretrained untuk ekstraksi embedding, menghasilkan 46.954 kemunculan CRIME_ACT. Pemilihan konfigurasi clustering dilakukan menggunakan kerangka multi-kriteria yang terdiri dari filter mega-cluster, filter cluster starved, dan Jensen-Shannon divergence antar distribusi LOC_TYPE per cluster sebagai kriteria perangkingan utama. Dari 21 konfigurasi yang diuji, no-aggregation pada K=18 terpilih sebagai pemenang dengan JS divergence sebesar 0.3598, menghasilkan 16 cluster layak yang mencakup pola seperti kekerasan seksual institusional, pelecehan seksual berbasis transit, penembakan di tempat komersial, dan KDRT tempat tinggal. Penelitian ini berkontribusi pada desain pipeline dua model yang memisahkan deteksi span dari representasi semantik, serta kerangka pemilihan clustering yang memprioritaskan utilitas downstream dibandingkan metrik internal konvensional.
============================================================
============================================================
=======
The rising crime trend in Indonesia and the limited access to incident-level official statistics motivate the need for analytical systems capable of automatically extracting crime event information from online news. This study develops an end-to-end pipeline that transforms Indonesian crime news narratives into a structured vulnerability map linking crime clusters to physical venue types, enriched with perpetrator type and target item context. A corpus of 13,937 articles was collected from detik.com between 1 January 2023 and 31 December 2024 and cleaned through Unicode normalization and boilerplate removal. A stratified sample of 400 articles was manually annotated using a four-entity BIO scheme: CRIME_ACT, LOC_TYPE, PERPETRATOR_TYPE, and TARGET_ITEM, at the character level. NER training was conducted in two phases: Phase 1 identified the optimal Indonesian pre-trained BERT checkpoints among four candidates (IndoBERT-IndoLEM, IndoBERT-IndoNLU, CahyaBERT, NusaBERT), with IndoBERT-IndoLEM as the best performer with a val F1 of 0.6830; Phase 2 applied Bayesian hyperparameter optimization via Optuna, raising val F1 to 0.7353 with a final test F1 of 0.6658. Dual-model inference was performed across the entire corpus, with the fine-tuned model handling span detection and the base pretrained model providing embedding extraction, producing 46,954 CRIME_ACT occurrences. The clustering configuration was selected using a multi-criterion framework consisting of a mega-cluster filter, a starved-cluster filter, and average pairwise Jensen-Shannon divergence between per-cluster LOC_TYPE distributions as the primary ranking criterion. Among 21 candidate configurations, no-aggregation at K=18 was selected as the winner with a JS divergence of 0.3598, producing 16 viable clusters that reveal patterns such as institutional sexual abuse, transit-based sexual harassment, commercial-premises shooting, and residential domestic violence. The study contributes a two-model pipeline design that separates span detection from semantic representation, and a clustering selection framework that prioritizes downstream task utility over conventional internal metrics.

Item Type: Thesis (Other)
Uncontrolled Keywords: Named Entity Recognition, Transformer, IndoBERT, berita kriminal, clustering semantik, pemetaan kerentanan, Jensen-Shannon divergence ========================================================== Named Entity Recognition, Transformer, BERT, crime news, semantic clustering, vulnerability mapping, Jensen-Shannon Divergence
Subjects: T Technology > T Technology (General) > T57.5 Data Processing
Divisions: Faculty of Information Technology > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: I Putu Raditya Partha Wijaya
Date Deposited: 27 Jul 2026 01:18
Last Modified: 27 Jul 2026 01:18
URI: http://repository.its.ac.id/id/eprint/137621

Actions (login required)

View Item View Item