Klasifikasi Aspek-Sentimen Multilabel Pada Ulasan Berbahasa Indonesia Aplikasi Duolingo Di Google Play Store Menggunakan IndoBERT

Putri, Aliffia Isma (2026) Klasifikasi Aspek-Sentimen Multilabel Pada Ulasan Berbahasa Indonesia Aplikasi Duolingo Di Google Play Store Menggunakan IndoBERT. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5026221130-Undergraduate_Thesis.pdf] Text
5026221130-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (8MB) | Request a copy

Abstract

Ulasan pengguna Duolingo di Google Play Store dapat memuat beberapa aspek dengan polaritas sentimen yang berbeda dalam satu teks. Kondisi tersebut menyebabkan analisis sentimen tingkat dokumen dengan satu label belum mampu merepresentasikan opini pengguna secara lengkap. Penelitian ini bertujuan mengevaluasi reliabilitas anotasi dan karakteristik dataset multilabel, membandingkan kinerja beberapa pendekatan klasifikasi aspek-sentimen, serta mengevaluasi kemampuan generalisasi model pada ulasan aplikasi pembelajaran bahasa lain. Pengumpulan data menghasilkan 67.603 ulasan Duolingo berbahasa Indonesia yang melalui proses penyaringan dan prapemrosesan awal sehingga diperoleh 55.103 ulasan bersih. Sebanyak 2.754 ulasan dipilih melalui proportional stratified sampling berdasarkan kategori panjang teks dan dianotasi secara independen oleh tiga anotator menggunakan sepuluh pasangan label aspek-sentimen. Reliabilitas anotasi diukur menggunakan Fleiss’ Kappa, sedangkan label akhir ditentukan melalui majority voting. Dataset ground truth dibagi menjadi data latih, validasi, dan uji dengan rasio 70:10:20. Pemodelan membandingkan lima skenario yang mencakup model klasik, deep learning, transformer, hybrid end-to-end, dan hybrid feature-based. Hasil pengukuran reliabilitas memperoleh rata-rata Fleiss’ Kappa sebesar 0,8862. Dataset memiliki label cardinality sebesar 1,1456, label density sebesar 0,1146, serta distribusi label yang tidak seimbang. Sebanyak 15 dari 45 pasangan label memiliki hubungan yang signifikan setelah koreksi Benjamini-Hochberg, tetapi seluruh kekuatan hubungannya berada pada kategori sangat lemah atau lemah. Pada evaluasi internal, IndoBERT dengan CNN 1D menghasilkan kinerja tertinggi dengan Micro F1-Score sebesar 0,8560, Micro-Precision sebesar 0,9018, Micro-Recall sebesar 0,8146, Hamming Loss sebesar 0,0314, dan Jaccard Score sebesar 0,7482. Analisis kesalahan menunjukkan 56 false positive dan 117 false negative, sehingga model masih lebih sering melewatkan label aktif aktual. Evaluasi generalisasi dilakukan tanpa pelatihan ulang terhadap 825 ulasan dari Busuu, Falou, dan LingoDeer. IndoBERT-Base memperoleh Micro F1-Score eksternal tertinggi sebesar 0,6164, diikuti IndoBERT dengan CNN 1D sebesar 0,6159. Seluruh model mengalami penurunan kinerja pada dataset eksternal. Hasil penelitian menunjukkan bahwa pendekatan hybrid end-to-end IndoBERT dengan CNN 1D efektif pada data internal, tetapi model dengan kinerja internal tertinggi belum tentu memiliki generalisasi terbaik ketika diterapkan pada distribusi data yang berbeda.
=====================================================================================================================================
Duolingo user reviews on the Google Play Store may discuss multiple aspects with different sentiment polarities within a single text. This characteristic means that document-level sentiment analysis using a single label cannot fully represent users’ opinions. This study aims to evaluate annotation reliability and multi-label dataset characteristics, compare several aspect-sentiment classification approaches, and assess model generalization on reviews from other language-learning applications. Data collection produced 67,603 Indonesian-language Duolingo reviews, which were filtered and initially preprocessed to obtain 55,103 clean reviews. A total of 2,754 reviews were selected through proportional stratified sampling based on text-length categories and independently annotated by three annotators using ten aspect-sentiment label pairs. Annotation reliability was measured using Fleiss’ Kappa, while the final labels were determined through majority voting. The resulting ground-truth dataset was divided into training, validation, and testing sets using a 70:10:20 ratio. Five modeling scenarios were compared, comprising classical machine learning, deep learning, transformer, end-to-end hybrid, and feature-based hybrid approaches. The annotation process achieved an average Fleiss’ Kappa of 0.8862. The dataset had a label cardinality of 1.1456, a label density of 0.1146, and an imbalanced label distribution. Fifteen of the 45 label pairs remained statistically significant after the Benjamini–Hochberg correction, although all associations were either very weak or weak. In the internal evaluation, IndoBERT with CNN 1D achieved the highest performance, with a Micro F1-Score of 0.8560, Micro-Precision of 0.9018, Micro-Recall of 0.8146, Hamming Loss of 0.0314, and Jaccard Score of 0.7482. Error analysis identified 56 false positives and 117 false negatives, indicating that the model was more likely to miss actual active labels. Generalization was evaluated without retraining on 825 reviews from Busuu, Falou, and LingoDeer. IndoBERT-Base achieved the highest external Micro F1-Score of 0.6164, closely followed by IndoBERT with CNN 1D at 0.6159. All models experienced performance degradation on the external dataset. These findings indicate that the end-to-end IndoBERT with CNN 1D approach is effective on internal data, but the model with the highest internal performance does not necessarily provide the strongest generalization under a different data distribution.

Item Type: Thesis (Other)
Uncontrolled Keywords: Generalisasi Model, IndoBERT, Klasifikasi Aspek-Sentimen, Klasifikasi Multilabel, Ulasan Duolingo, Aspect-Sentiment Classification, Duolingo Reviews, IndoBERT, Model Generalization, Multi-Label Classification
Subjects: T Technology > T Technology (General) > T57.5 Data Processing
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 57201-(S1) Undergraduate Thesis
Depositing User: Aliffia Isma Putri
Date Deposited: 03 Aug 2026 04:41
Last Modified: 03 Aug 2026 04:41
URI: http://repository.its.ac.id/id/eprint/142228

Actions (login required)

View Item View Item