Yaqutah, Athaya Rohadatul (2026) Pseudolabelling Data Emosi Kesehatan Mental Menggunakan Semi-Supervised Learning dan Contrastive Learning. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221235-Undergraduate_Thesis.pdf Restricted to Repository staff only Download (20MB) | Request a copy |
Abstract
Klasifikasi emosi kesehatan mental pada teks media sosial berbahasa Indonesia menghadapi tantangan utama berupa keterbatasan data berlabel berkualitas tinggi yang memerlukan keahlian khusus dan biaya tinggi dalam proses anotasinya. Penelitian ini mengembangkan kerangka semi-supervised learning dua fase yang mengintegrasikan contrastive learning dan pseudo-labeling untuk klasifikasi multi-label delapan emosi kesehatan mental (marah, sedih, merasa tidak berharga, niat bunuh diri, kosong, kesepian, gangguan kognitif, dan putus asa) pada teks media sosial berbahasa Indonesia. Dataset penelitian terdiri dari 15.369 postingan dari platform Twitter, Reddit, Quora, dan Threads, dengan 7.000 data berlabel yang dihasilkan melalui pelabelan otomatis menggunakan Google Gemini API dan 8.369 data tanpa label. Kerangka yang diusulkan mengimplementasikan empat skenario pengujian: (1) baseline yang mengintegrasikan tiga strategi contrastive learning (BAL, SCL, JSCL) dengan pseudo-labeling menggunakan parameter default; (2) hyperparameter tuning pada metode terbaik menggunakan Optuna; (3) iterative pseudo-labeling dengan filter minoritas dan BCE pos_weight; (4) retraining menggunakan dataset hasil augmentasi dan pseudo-label. Evaluasi menggunakan metrik F1-Score (Micro/Macro), Precision, dan Recall. Hasil terbaik diperoleh setelah tuning hiperparameter dan penyeimbangan bobot kelas (pos_weight), dengan kombinasi IndoBERTweet dan BAL menghasilkan F1-Macro 0,7491 dan F1-Micro 0,7506. Pada tahap retraining dengan dataset augmentasi pseudo-label dari korpus Reddit, kombinasi Reddit-BAL mencapai F1-Macro 0,7454, F1-Micro 0,7496, dan Precision Macro 0,7505. Penelitian ini menunjukkan bahwa hyperparameter tuning dan pos_weight efektif meningkatkan performa klasifikasi, dengan BAL terbukti paling stabil dibandingkan SCL dan JSCL. Augmentasi pseudo-label pada tahap retraining kompetitif namun belum melampaui tahap tuning, menunjukkan optimasi hyperparameter memberikan kontribusi lebih besar dibandingkan penambahan data augmentasi.
=====================================================================================================================================
The classification of mental health emotions in Indonesian social media texts faces a major challenge due to the limited availability of high-quality labeled data, which requires specialized expertise and high costs in the annotation process. This research develops a two-phase semi-supervised learning framework that integrates contrastive learning and pseudo-labeling for multi-label classification of eight mental health emotions (anger, sadness, worthlessness, suicide intent, emptiness, loneliness, cognitive dysfunction, and hopelessness) in Indonesian social media texts. The research dataset consists of 15,369 posts from Twitter, Reddit, Quora, and Threads platforms, with 7,000 labeled data generated through automatic labeling using Google Gemini API and 8,369 unlabeled data. The proposed framework implements four testing scenarios: (1) baseline integrating three contrastive learning strategies (BAL, SCL, JSCL) with pseudo-labeling using default parameters; (2) hyperparameter tuning on the best method using Optuna; (3) iterative pseudo-labeling with minority filtering and BCE pos_weight; (4) retraining using augmented and pseudo-labeled datasets. Evaluation employs F1-Score (Micro/Macro), Precision, and Recall metrics. The best results were achieved after hyperparameter tuning and class weighting (pos_weight), with IndoBERTweet-BAL yielding F1-Macro 0.7491 and F1-Micro 0.7506. In the retraining stage using a Reddit-based pseudo-labeled augmented dataset, Reddit-BAL achieved F1-Macro 0.7454, F1-Micro 0.7496, and Precision Macro 0.7505. This research shows that hyperparameter tuning and pos_weight effectively improve classification performance, with BAL proving most stable compared to SCL and JSCL. Pseudo-label-based augmentation in the retraining stage was competitive but did not surpass the tuning stage, indicating hyperparameter optimization contributed more than added augmented data.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Semi-Supervised Learning, Contrastive Learning, Pseudo-Labeling, Klasifikasi Multi-Label, Emosi Kesehatan Mental Semi-Supervised Learning, Contrastive Learning, Pseudo-Labeling, Multi-Label Classification, Mental Health Emotion |
| Subjects: | T Technology > T Technology (General) T Technology > T Technology (General) > T174 Technological forecasting T Technology > T Technology (General) > T57.5 Data Processing T Technology > T Technology (General) > T58.62 Decision support systems |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Athaya Rohadatul Yaqutah |
| Date Deposited: | 24 Jul 2026 04:28 |
| Last Modified: | 24 Jul 2026 04:29 |
| URI: | http://repository.its.ac.id/id/eprint/136886 |
Actions (login required)
![]() |
View Item |
