Klasifikasi Multilabel Emosi Menggunakan Pendekatan Iterative Pseudo-Labelling dan Model Transformer

Putri, Nadya Saraswati (2026) Klasifikasi Multilabel Emosi Menggunakan Pendekatan Iterative Pseudo-Labelling dan Model Transformer. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221246-Undergraduate_Thesis.pdf] Text
5025221246-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (11MB) | Request a copy

Abstract

Kesehatan mental merupakan isu yang semakin banyak dibahas di media sosial. Namun, deteksi otomatis emosi terkait kesehatan mental pada teks berbahasa Indonesia masih menghadapi tantangan, terutama karena sifat multilabel, yaitu satu unggahan dapat mengandung lebih dari satu emosi secara bersamaan, serta adanya ketidakseimbangan distribusi label. Penelitian ini membandingkan empat metode klasifikasi multilabel, yaitu Binary Relevance, Classifier Chains, Multi-Task Learning, dan Label Attention Mechanism, untuk mengidentifikasi delapan label emosi kesehatan mental pada teks media sosial berbahasa Indonesia. Data penelitian dikumpulkan dari Reddit, Quora, Threads, dan Twitter/X. Dari 17.018 unggahan yang diperoleh, sebanyak 15.369 data dinyatakan layak setelah proses penyaringan. Sebanyak 7.000 data kemudian dianotasi dengan bantuan Gemini AI dan dibagi menggunakan metode multi-label stratified splitting. Keempat metode klasifikasi diimplementasikan pada empat model transformer, yaitu MentalBERT, IndoBERTweet, IndoJavE-IndoBERTweet, dan XLM-RoBERTa, dengan optimasi hiperparameter menggunakan Optuna. Sebelum dilakukan penyeimbangan data, metode Classifier Chains dengan model XLM-RoBERTa menunjukkan kinerja terbaik dengan nilai F1-Micro sebesar 0,7830 dan F1-Macro sebesar 0,7802. Untuk mengatasi ketidakseimbangan label, penelitian ini menerapkan iterative pseudo-labeling berbasis self-training yang dikombinasikan dengan augmentasi data melalui back-translation. Pendekatan tersebut meningkatkan jumlah data berlabel sebesar 36,9% dan menaikkan nilai F1-Macro validasi dari 0,6182 menjadi 0,7360 setelah lima putaran self-training. Setelah proses penyeimbangan data, kombinasi Classifier Chains dan XLM-RoBERTa tetap menghasilkan performa terbaik, dengan nilai F1-Micro sebesar 0,7870, F1-Macro sebesar 0,7846, Precision sebesar 0,7602, dan Recall sebesar 0,8158. Hasil analisis menunjukkan bahwa model lebih mampu mengenali emosi yang dinyatakan secara eksplisit, sementara emosi yang bersifat implisit atau memiliki tumpang tindih makna antarlabel masih sulit dibedakan. Temuan ini menunjukkan bahwa klasifikasi multilabel menggunakan Classifier Chains, yang didukung strategi penyeimbangan data berbasis pseudo-labeling, efektif dalam meningkatkan kinerja klasifikasi multilabel emosi kesehatan mental pada teks media sosial berbahasa Indonesia.
====================================================================================================================================
Mental health has become an increasingly discussed issue on social media. However, the automatic detection of mental health-related emotions in Indonesian-language text remains challenging, particularly due to the multi-label nature of the task, in which a single post may contain more than one emotion simultaneously, as well as the imbalanced distribution of emotion labels. This study compares four multi-label classification methods, namely Binary Relevance, Classifier Chains, Multi-Task Learning, and Label Attention Mechanism, to identify eight mental health emotion labels in Indonesian social media texts. The dataset was collected from Reddit, Quora, Threads, and Twitter/X. Of the 17,018 posts gathered, 15,369 were retained after the data filtering process. A total of 7,000 posts were then annotated with the assistance of Gemini AI and partitioned using multi-label stratified splitting. The four classification methods were implemented using four transformer-based models: MentalBERT, IndoBERTweet, IndoJavE-IndoBERTweet, and XLM-RoBERTa, with hyperparameter optimization performed using Optuna. Before data balancing, Classifier Chains with XLM-RoBERTa achieved the best performance, obtaining an F1-Micro score of 0.7830 and an F1-Macro score of 0.7802. To address label imbalance, this study employed iterative pseudo-labeling based on self-training and combined it with data augmentation through back-translation. This approach increased the amount of labeled data by 36.9% and improved the validation F1-Macro score from 0.6182 to 0.7360 after five rounds of self-training. After data balancing, the combination of Classifier Chains and XLM-RoBERTa remained the best-performing configuration, achieving an F1-Micro score of 0.7870, an F1-Macro score of 0.7846, Precision of 0.7602, and Recall of 0.8158. The analysis showed that the model was more effective at identifying explicitly expressed emotions, while implicitly expressed emotions and labels with overlapping meanings remained difficult to distinguish. These findings indicate that multi-label classification using Classifier Chains, supported by a pseudo-labeling-based data balancing strategy, is effective in improving the performance of mental health emotion classification in Indonesian social media texts.

Item Type: Thesis (Other)
Uncontrolled Keywords: Klasifikasi Multilabel, Emosi Kesehatan Mental, Transformer, Iterative Pseudo-Labeling, Supervised Learning Multi-Label Classification, Mental Health Emotion, Transformer, Iterative Pseudo-Labeling, Supervised Learning
Subjects: Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Nadya Saraswati Putri
Date Deposited: 25 Jul 2026 07:55
Last Modified: 25 Jul 2026 07:55
URI: http://repository.its.ac.id/id/eprint/137855

Actions (login required)

View Item View Item