CC : Virtual-Youtuber (Vtuber) Berbasis Kecerdasan Buatan Dengan Pengaplikasian Large Language Model (LLM) Dengan Retrieval-Augmented Generation (RAG)

Shaula, Shaula Aljauhara Riyadi (2026) CC : Virtual-Youtuber (Vtuber) Berbasis Kecerdasan Buatan Dengan Pengaplikasian Large Language Model (LLM) Dengan Retrieval-Augmented Generation (RAG). Other thesis, Institut Teknologi Sepuluh nopember.

[thumbnail of 5025201265_Undergraduate_Thesis.pdf] Text
5025201265_Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (8MB) | Request a copy

Abstract

Industri hiburan digital, khususnya Vtuber, berkembang pesat dengan hadirnya entitas AI seperti Neuro-sama yang populer sebagai Vtuber berbasis AI. Namun, VTuber konvensional masih bergantung pada operator manusia sehingga esensi "virtual"-nya belum sepenuhnya tercapai. Di sisi lain, penggunaan Large Language Model (LLM) sebagai otak digital rentan mengalami halusinasi informasi dan keterbatasan memori kontekstual. Penelitian ini mengembangkan agen VTuber AI bernama "CC" yang mengintegrasikan LLM dengan Retrieval-Augmented Generation (RAG) sebagai basis pengetahuan untuk menekan halusinasi. Fokusnya adalah menciptakan penyiar virtual berkarakter kasual dan sarkastik yang interaktif namun ringan dijalankan pada perangkat terbatas. Eksperimen membandingkan arsitektur Causal LM dan Sequence-to-Sequence LM menggunakan dua dataset (CSV hasil terjemahan bahasa Inggris dan JSONL berbahasa Indonesia asli), yang dilatih dengan pendekatan Low-Rank Adaptation (LoRA) dan tiga rekayasa parameter (learning rate dan LoRA memory) untuk mengetahui efek perubahan hyperparameter terhadap model. Hasil pengujian menunjukkan model Causal LM berbasis Gemma3 paling unggul untuk tugas roleplay, dengan skor kedekatan semantik IndoBERT 0,67/1, G-Eval Character 4,0/5, dan Context 4,2/5. Sistem RAG berhasil menekan halusinasi dengan skor faithfulness 0,66/1 dan answer relevance 0,57/1. Kesimpulannya, agen VTuber AI "CC" berhasil diintegrasikan ke sistem backend lokal secara stabil dan membuktikan kelayakan LLM sebagai inovasi konten penyiaran di industri hiburan digital.
===================================================================================================================================
The digital entertainment industry, particularly Vtubers, has grown rapidly with the emergence of AI entities such as Neuro-sama, a popular AI-based Vtuber. However, conventional VTubers still rely on human operators, meaning their "virtual" essence has not been fully realized. Meanwhile, the use of Large Language Models (LLMs) as a digital brain is prone to information hallucination and limited contextual memory. This research develops an AI VTuber agent named "CC" that integrates an LLM with Retrieval-Augmented Generation (RAG) as a knowledge base to reduce hallucination. The main focus is to create an interactive virtual broadcaster with a casual and sarcastic character, while remaining lightweight enough to run on limited devices. Experiments compared Causal LM and Sequence-to-Sequence LM architectures using two datasets (a CSV translated from English and a JSONL originally in Indonesian), trained using the Low-Rank Adaptation (LoRA) approach with three parameter configurations (learning rate and LoRA memory) to find the effect of hyperparameter change toward model. Test results show that the Gemma3-based Causal LM model performed best for roleplay tasks, achieving an IndoBERT semantic similarity score of 0.67/1, a G-Eval Character score of 4.0/5, and a Context score of 4.2/5. The RAG system successfully reduced hallucination, with a faithfulness score of 0.66/1 and an answer relevance score of 0.57/1. In conclusion, the "CC" AI VTuber agent was successfully and stably integrated into a local backend system, demonstrating the feasibility of LLMs as an innovation in broadcasting content within the digital entertainment industry.

Item Type: Thesis (Other)
Uncontrolled Keywords: VTuber AI, Large Language Model, Causal LM, Sequence 2 Sequence LM, Retrieval-Augmented Generation, Gemma3, Low Rank Adaptation.
Subjects: T Technology > T Technology (General) > T11 Technical writing. Scientific Writing
T Technology > T Technology (General) > T385 Visualization--Technique
T Technology > T Technology (General) > T59.7 Human-machine systems.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Shaula Aljauhara Riyadi
Date Deposited: 05 Aug 2026 04:25
Last Modified: 05 Aug 2026 04:25
URI: http://repository.its.ac.id/id/eprint/143937

Actions (login required)

View Item View Item