Syahidan, Danial Rajiv and Najla, Adinda Shafa and Aji, Panji Seno and I'tishom, Muhammad Khibban (2026) Perlindungan Anak di Ruang Digital (PAD) CPT + SFT - LLM Program AI Talent Factory (AITF). Project Report. [s.n.], [s.l.]. (Unpublished)
|
Text
5054231004_5053231004_5054231023_5025231126-Project_Report.pdf - Accepted Version Download (2MB) |
Abstract
Perkembangan media sosial yang sangat pesat meningkatkan risiko paparan konten digital yang tidak sesuai bagi anak dan remaja. Berbagai bentuk konten seperti perjudian daring, kekerasan, ujaran kebencian, pelecehan, eksploitasi seksual, maupun perilaku menyakiti diri sendiri tersebar dengan cepat melalui berbagai platform digital. Kondisi tersebut menjadi tantangan bagi Kementerian Komunikasi dan Digital (KOMDIGI) dalam melakukan moderasi konten secara efektif, mengingat tingginya volume data yang harus diproses serta keterbatasan metode moderasi konvensional yang masih bergantung pada pencocokan kata kunci dan verifikasi manual.
Melalui program Artificial Intelligence Talent Factory (AITF), Tim PAD 1 Institut Teknologi Sepuluh Nopember mengembangkan model Large Language Model (LLM) yang dikhususkan untuk melakukan klasifikasi rating usia terhadap konten teks berbahasa Indonesia dalam rangka mendukung implementasi Perlindungan Anak di Ruang Digital (PAD). Model dikembangkan melalui dua tahapan utama, yaitu Continued Pre-Training (CPT) menggunakan korpus teks spesifik domain perlindungan anak dan Supervised Fine-Tuning (SFT) menggunakan dataset instruksi yang telah dikurasi agar model mampu menghasilkan klasifikasi rating usia beserta penjelasan secara konsisten. Hasil pelatihan kemudian diintegrasikan menjadi layanan inferensi berbasis REST API sehingga dapat digunakan sebagai komponen pada sistem moderasi konten digital.
Hasil pengujian menunjukkan bahwa pendekatan adaptasi domain melalui CPT yang dilanjutkan dengan SFT mampu meningkatkan kemampuan model dalam memahami konteks bahasa Indonesia, termasuk penggunaan bahasa tidak baku, istilah media sosial, serta konten sensitif yang berkaitan dengan perlindungan anak. Model yang dihasilkan mampu mengklasifikasikan konten ke dalam kategori Semua Umur (SU), 7+, 13+, 15+, 18+, maupun Konten Terlarang, disertai penjelasan yang mendukung proses verifikasi oleh moderator. Seluruh komponen sistem berhasil diimplementasikan dan dideploy sebagai Minimum Viable Product (MVP) yang siap diintegrasikan pada ekosistem Perlindungan Anak di Ruang Digital KOMDIGI.
=====================================================================================================================================
The rapid growth of social media has significantly increased the risk of children and adolescents being exposed to inappropriate digital content. Various forms of harmful content, including online gambling, violence, hate speech, harassment, sexual exploitation, and self-harm-related material, spread rapidly across digital platforms. This poses a major challenge for the Ministry of Communication and Digital Affairs (KOMDIGI) in conducting effective content moderation, given the massive volume of data that must be processed and the limitations of conventional moderation methods that primarily rely on keyword matching and manual verification.
Through the Artificial Intelligence Talent Factory (AITF) program, Team PAD 1 from Institut Teknologi Sepuluh Nopember developed a Large Language Model (LLM) specifically designed to classify age ratings for Indonesian-language textual content in support of the Child Protection in Digital Spaces (Perlindungan Anak di Ruang Digital, PAD) initiative. The model was developed through two main stages: Continued Pre-Training (CPT) using a domain-specific corpus related to child protection, followed by Supervised Fine-Tuning (SFT) on a curated instruction dataset to enable the model to consistently generate age-rating classifications along with explanatory reasoning. The trained model was subsequently integrated into a REST API-based inference service, allowing it to function as a component of a digital content moderation system.
Experimental results demonstrate that the domain adaptation approach through CPT followed by SFT effectively improves the model's ability to understand the context of the Indonesian language, including informal expressions, social media terminology, and sensitive content related to child protection. The resulting model is capable of classifying content into the categories General Audience (SU), 7+, 13+, 15+, 18+, and Prohibited Content, while providing explanations that assist moderators during the verification process. All system components were successfully implemented and deployed as a Minimum Viable Product (MVP), making the solution ready for integration into KOMDIGI's Child Protection in Digital Spaces ecosystem.
| Item Type: | Monograph (Project Report) |
|---|---|
| Uncontrolled Keywords: | Perlindungan Anak di Ruang Digital, Child Protection in Digital Spaces, Large Language Model, Large Language Model, Continued Pre-Training, Continued Pre-Training, Supervised Fine-Tuning, Supervised Fine-Tuning, Moderasi Konten, Content Moderation, Klasifikasi Rating Usia, Age Rating Classification. |
| Subjects: | T Technology > T Technology (General) > T57.5 Data Processing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Panji Seno Aji |
| Date Deposited: | 12 Jul 2026 08:14 |
| Last Modified: | 12 Jul 2026 08:14 |
| URI: | http://repository.its.ac.id/id/eprint/134690 |
Actions (login required)
![]() |
View Item |
