Adi, Gandrung Ghafar Hernendra (2026) Disinformasi Fitnah Ujaran Kebencian - CPT & SFT. Project Report. [s.n.]. (Unpublished)
|
Text
5054231008-Project_Report.pdf - Accepted Version Restricted to Repository staff only Download (1MB) | Request a copy |
Abstract
Kerja Praktik ini berfokus pada pengembangan dan adaptasi Large Language Model (LLM) yang terspesialisasi untuk tugas deteksi konten negatif berbasis teks, meliputi Disinformasi, Fitnah, dan Ujaran Kebencian (DFK) pada media sosial berbahasa Indonesia. Pengembangan model dilakukan melalui dua tahapan utama, yaitu Continued Pre-Training (CPT) dan Supervised Fine-Tuning (SFT). Tahap CPT menggunakan korpus spesifik domain DFK sebanyak 374.930 baris data (227,4 juta token) untuk memperkaya pemahaman model dasar. Tahap SFT dilakukan untuk melatih model agar mampu mengklasifikasikan konten dan menghasilkan penalaran (reasoning) yang relevan. Eksperimen dilakukan dengan menggunakan arsitektur dasar Ministral-3-8B-Base-2512. Hasil pengujian menunjukkan bahwa tahap CPT berhasil menurunkan nilai perplexity pada domain target dari 7.7082 menjadi 5.8837 (peningkatan adaptasi 23,67%). Pada tahap SFT, model terbaik mencapai nilai akurasi dan F1-score sebesar 0,99 pada data uji, serta nilai BERTScore-F1 hingga 0,82 yang menunjukkan kualitas keselarasan penalaran dengan jawaban acuan. Model final ini kemudian di-deploy dalam bentuk layanan Application Programming Interface (API) menggunakan platform komputasi serverless, yang diintegrasikan ke dalam purwarupa Minimum Viable Product (MVP) untuk mendukung pengawasan ruang digital oleh Kementerian Komunikasi dan Digital RI.
===================================================================================================================================
This internship project focuses on the development and adaptation of a specialized Large Language Model (LLM) for text-based negative content detection, encompassing Disinformation, Slander, and Hate Speech (DFK) on Indonesian social media platforms. The model development is executed through two main stages: Continued Pre-Training (CPT) and Supervised Fine-Tuning (SFT). The CPT stage utilizes a domain-specific DFK corpus consisting of 374,930 rows of data (227.4 million tokens) to enrich the foundational model's understanding. The SFT stage is conducted to train the model to accurately classify content and generate relevant reasoning. Experiments are carried out using the Ministral-3-8B-Base-2512 base architecture. The evaluation results indicate that the CPT stage successfully reduces the perplexity value on the target domain from 7.7082 to 5.8837, representing a 23.67% improvement in domain adaptation. In the SFT stage, the best-performing model achieves an accuracy and F1-score of 0.99 on the test dataset, along with a BERTScore-F1 of up to 0.82, demonstrating high alignment quality between the generated reasoning and the reference answer. Finally, the definitive model is deployed as an Application Programming Interface (API) service using a serverless computing platform, which is integrated into a Minimum Viable Product (MVP) prototype to support digital space surveillance by the Ministry of Communication and Digital Affairs of the Republic of Indonesia.
| Item Type: | Monograph (Project Report) |
|---|---|
| Uncontrolled Keywords: | Continued Pre-Training, Disinformasi, Fitnah, Large Language Model, Supervised Fine-Tuning, Ujaran Kebencian, Disinformation, Slander, Hate Speech |
| Subjects: | T Technology > T Technology (General) T Technology > T Technology (General) > T57.5 Data Processing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Artificial Intelligence Engineering > 55283-(S1) Undergraduate Thesis |
| Depositing User: | Gandrung Gafar |
| Date Deposited: | 15 Jul 2026 06:08 |
| Last Modified: | 15 Jul 2026 06:08 |
| URI: | http://repository.its.ac.id/id/eprint/134834 |
Actions (login required)
![]() |
View Item |
