Fadila, Chayun (2026) Analisis Sentimen Berbasis Aspek pada Komentar dan Emotikon Live Streaming TikTok dengan Pendekatan BERT dan LLM. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6026242011-Master_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (15MB) | Request a copy |
Abstract
Live streaming TikTok telah menjadi ruang interaksi real-time antara kreator dan penonton. Komentar yang muncul selama siaran memberikan umpan balik instan terhadap konten yang disiarkan. Pemahaman terhadap sentimen pada komentar tersebut penting untuk mengukur engagement penonton dan persepsi terhadap konten live. Penelitian ini bertujuan merancang, mengimplementasikan, dan membandingkan pendekatan BERT dan Large Language Model (LLM) untuk Aspect-Based Sentiment Analysis (ABSA) pada komentar dan emote live streaming TikTok berbahasa Indonesia. Tiga pendekatan dikembangkan. Pertama, IndoBERT diadaptasi dengan fine-tuning Low-Rank Adaptation (LoRA) sebagai model deep learning, menghasilkan arsitektur dual-head yang mengklasifikasikan aspek dan emosi secara simultan. Kedua, ChatGPT (gpt-4o-mini) digunakan dengan teknik few-shot prompting sebagai sistem auto-labeling untuk membangun dataset beranotasi. Ketiga, Llama-3.1-8B-Instruct diadaptasi melalui fine-tuning LoRA sebagai pendekatan LLM open-source untuk pelabelan dan klasifikasi. Dataset terdiri dari komentar live streaming TikTok berbahasa Indonesia yang mencakup teks dan emote, dengan anotasi empat kategori aspek (Host, Konten, Interaksi, Topik), delapan kelas emosi, dan tiga polaritas sentimen. Data dikumpulkan melalui scraping pada sesi live streaming yang sedang berlangsung dengan memperhatikan kode etik dan privasi kreator. Praproses data meliputi pembersihan teks, normalisasi kosakata informal, konversi emote ke token semantik, dan penyeimbangan kelas melalui oversampling. Evaluasi performa model menggunakan metrik akurasi, presisi, recall, dan F1-score. Hasil penelitian menunjukkan IndoBERT+LoRA mencapai akurasi 98,37% (F1-makro 0,9836) pada klasifikasi aspek dan 97,21% (F1-makro 0,9507) pada klasifikasi emosi. Capaian ini jauh mengungguli Llama+LoRA yang hanya mencapai akurasi 55,59% pada aspek dan 52,07% pada emosi. Di sisi lain, ChatGPT terbukti unggul sebagai sistem auto-labeling, dengan keberhasilan melabeli seluruh 23.958 komentar tanpa satu pun kegagalan parsing. Temuan ini menunjukkan bahwa model BERT yang diadaptasi secara efisien melalui LoRA mampu mengungguli LLM generatif pada tugas klasifikasi berbahasa Indonesia, meskipun performa pada kelas-kelas dengan sampel sangat minoritas perlu dimaknai secara hati-hati mengingat skema oversampling yang diterapkan sebelum pembagian data. Penelitian ini berkontribusi terhadap pengembangan metode ABSA pada live streaming TikTok berbahasa Indonesia, khususnya dalam menangani komentar yang singkat, spontan, dan kaya emotikon. Hasil penelitian ini dapat dimanfaatkan sebagai dasar optimasi strategi dan pengembangan konten live streaming yang lebih efektif.
====================================================================================================================================
TikTok live streaming has become a real-time interaction space between creators and viewers. Comments that appear during a broadcast provide instant feedback on the content being streamed. Understanding the sentiment embedded in these comments is essential for measuring viewer engagement and perception of live content. This study aims to design, implement, and compare BERT and Large Language Model (LLM) approaches for Aspect-Based Sentiment Analysis (ABSA) on Indonesian-language comments and emotes from TikTok live streams. Three approaches were developed. First, IndoBERT was adapted through Low-Rank Adaptation (LoRA) fine-tuning as the deep learning model, resulting in a dual-head architecture that classifies aspect and emotion simultaneously. Second, ChatGPT (gpt-4o-mini) was employed with a few-shot prompting technique as an auto-labeling system to construct an annotated dataset. Third, Llama-3.1-8B-Instruct was adapted through LoRA fine-tuning as an open-source LLM approach for labeling and classification. The dataset consisted of Indonesian-language TikTok live streaming comments encompassing both text and emotes, annotated with four aspect categories (Host, Content, Interaction, Topic), eight emotion classes, and three sentiment polarities. Data were collected through scraping of ongoing live streaming sessions with due consideration for ethical guidelines and creator privacy. Data preprocessing included text cleaning, informal vocabulary normalization, emote-to-semantic-token conversion, and class balancing through oversampling. Model performance was evaluated using accuracy, precision, recall, and F1-score metrics. The results show that IndoBERT+LoRA achieved an accuracy of 98.37% (macro F1-score of 0.9836) on aspect classification and 97.21% (macro F1-score of 0.9507) on emotion classification. This performance far surpassed that of Llama+LoRA, which achieved only 55.59% accuracy on aspect classification and 52.07% on emotion classification. Meanwhile, ChatGPT proved superior as an auto-labeling system, successfully labeling all 23,958 comments without a single parsing failure. These findings indicate that a BERT model efficiently adapted through LoRA is able to outperform a generative LLM on Indonesian-language classification tasks, although performance on extremely minority classes should be interpreted with caution given the oversampling scheme applied prior to data splitting. This study contributes to the development of ABSA methods for Indonesian-language TikTok live streaming, particularly in handling comments that are short, spontaneous, and rich in emoticons. The findings of this study may serve as a foundation for optimizing strategy and developing more effective live streaming content.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | Aspect-Based Sentiment Analysis, Sentiment Analisis, Emote, Live Streaming, Tiktok, Large Language Model, Deep Learning, BERT, ChatGPT, Llama, LoRA |
| Subjects: | T Technology > T Technology (General) > T57.8 Nonlinear programming. Support vector machine. Wavelets. Hidden Markov models. T Technology > TA Engineering (General). Civil engineering (General) T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7882.P3 Pattern recognition systems |
| Divisions: | Faculty of Information and Communication Technology > Information Systems > 59101-(S2) Master Thesis |
| Depositing User: | Chayun Fadila |
| Date Deposited: | 29 Jul 2026 01:48 |
| Last Modified: | 29 Jul 2026 01:48 |
| URI: | http://repository.its.ac.id/id/eprint/139195 |
Actions (login required)
![]() |
View Item |
