Putra, Alief Gilang Permana (2026) Asesmen Kepribadian Pada Video Singkat Menggunakan Pendekatan Hibrida: Prediksi Big Five Personality Berbasis Transformer dan Interpretasi Deskriptif Berbasis LLM. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221193-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (41MB) | Request a copy |
Abstract
Penilaian kepribadian tradisional berbasis kuesioner sering kali rentan terhadap bias subjektivitas dan pemalsuan respons. Meskipun teknologi kecerdasan buatan telah memungkinkan prediksi kepribadian secara otomatis berdasarkan data visual, hasil keluaran yang diberikan umumnya masih berupa skor numerik kontinu yang sulit dipahami secara kualitatif oleh pengguna awam. Penelitian ini mengusulkan sistem asesmen kepribadian hibrida pada potongan video singkat dengan menggabungkan model regresi Transformer untuk memprediksi skor Big Five Personality dan Large Language Model (LLM) untuk menghasilkan interpretasi profil kepribadian yang deskriptif. Metodologi penelitian mencakup ekstraksi frame gambar dari dataset ChaLearn First Impression, yang kemudian diproses menggunakan berbagai varian model Transformer (ViT, SwinV2, PVTv2) dengan modifikasi custom regression head yang terdiri dari lapisan linear, fungsi aktivasi GELU, dropout, dan sigmoid. Hasil prediksi regresi selanjutnya diintegrasikan ke dalam LLM melalui teknik prompt engineering multimodal dengan menerapkan skenario standard prompting dan roleplay prompting. Hasil eksperimen menunjukkan bahwa arsitektur ViT-B/16 AugReg (IN21k) dengan penerapan skenario arch tuning beserta augmentasi merupakan model prediksi regresi terbaik, dengan tingkat akurasi mencapai 0,911 dan nilai coefficient of determination (R²) sebesar 0,416. Pada tahap interpretasi teks yang dievaluasi menggunakan metrik G-Eval, penelitian ini mengimplementasikan pendekatan dua model alternatif. Model Qwen3-VL-32B-Instruct ditetapkan sebagai LLM utama yang meraih skor keseluruhan tertinggi (0,9136) dengan narasi persona psikolog yang kuat, sedangkan model Gemma-4-31B-IT ditetapkan sebagai alternatif karena menawarkan keseimbangan kinerja yang lebih baik, dengan keunggulan pada stabilitas akurasi psikologis (0,9504) serta kedisiplinan format output. Evaluasi tambahan oleh psikolog profesional terhadap kedua model tersebut menunjukkan skor rata-rata sebesar 52–61% dari skor maksimal pada empat aspek penilaian, mengindikasikan bahwa kualitas interpretasi yang dihasilkan LLM belum sepenuhnya setara dengan standar penilaian psikolog profesional, meskipun telah dinilai tinggi oleh G-Eval.
==================================================================================================================================
Traditional questionnaire-based personality assessments are often prone to subjectivity and response falsification. Although artificial intelligence technology has enabled automated personality predictions based on visual data, the output is generally still in the form of continuous numerical scores that are difficult for common users to interpret qualitatively. This study proposes a hybrid personality assessment system for short video clips by combining a Transformer regression model to predict Big Five Personality scores and a Large Language Model (LLM) to generate descriptive interpretations of personality profiles. The research methodology involves extracting image frames from the ChaLearn First Impression dataset, which are then processed using various variants of the Transformer model (ViT, SwinV2, PVTv2) with a custom regression head consisting of a linear layer, the GELU activation function, dropout, and a sigmoid function. The regression prediction results were then integrated into the LLM via multimodal prompt engineering techniques, applying both standard prompting and roleplay prompting scenarios. The experimental results show that the ViT-B/16 AugReg (IN21k) architecture, with the implementation of arch tuning and data augmentation, is the best regression prediction model, achieving an accuracy of 0.911 and a coefficient of determination (R²) of 0.416. In the text interpretation stage, which was evaluated using the G-Eval metric, this study implemented two alternative model approaches. The Qwen3-VL-32B-Instruct model was identified as the primary LLM, achieving the highest overall score (0.9136) with a strong psychologist persona narrative, while the Gemma-4-31B-IT model was identified as the alternative due to its better performance balance, excelling in psychological accuracy stability (0.9504) and output format consistency. An additional evaluation by a professional psychologist on both models yielded average scores of only 52–61% of the maximum score across four assessment aspects, indicating that the interpretive quality generated by the LLMs does not yet fully match the standard of a professional psychologist, despite receiving high scores from G-Eval.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Asesmen Kepribadian, Big Five Personality, ChaLearn, Large Language Model, Transformer |
| Subjects: | Q Science > QA Mathematics > QA336 Artificial Intelligence Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science) |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Alief Gilang Permana Putra |
| Date Deposited: | 21 Jul 2026 09:30 |
| Last Modified: | 21 Jul 2026 09:30 |
| URI: | http://repository.its.ac.id/id/eprint/136012 |
Actions (login required)
![]() |
View Item |
