Estimasi Kepribadian Berdasarkan Analisis Data Suara Menggunakan Pendekatan Deep Learning Dengan Emosi Sebagai Fitur Antara

Syuhada, Fathurazka Gamma (2026) Estimasi Kepribadian Berdasarkan Analisis Data Suara Menggunakan Pendekatan Deep Learning Dengan Emosi Sebagai Fitur Antara. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221128-Undergraduate_Thesis.pdf] Text
5025221128-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Suara manusia mengandung informasi nonverbal yang kaya, termasuk emosi dan kepribadian. Penilaian kepribadian secara konvensional melalui kuesioner bersifat subjektif,
sehingga dibutuhkan pendekatan otomatis dan objektif berbasis suara. Penelitian ini mengembangkan pendekatan dua tahap berbasis deep learning untuk mengestimasi kepribadian dari data suara, dengan memanfaatkan deteksi emosi sebagai fitur antara. Namun, deteksi emosi dari sinyal suara memiliki tantangan tersendiri karena pola emosi tidak hanya ditentukan oleh karakteristik spektral lokal seperti frekuensi, harmonik dan intensitas, tetapi juga oleh dinamika temporal sinyal suara sepanjang ucapan, sehingga model yang hanya menangkap salah satu aspek tersebut cenderung kurang optimal. Untuk mengatasi tantangan tersebut, pada tahap pertama dikembangkan model CNN-BiLSTM, di mana CNN mengekstraksi pola spektral local dan BiLSTM memodelkan dependensi temporal dari dua arah, untuk mendeteksi delapan kelas emosi dari mel-spectrogram menggunakan dataset RAVDESS, yang mencapai akurasi 79,17% pada testing set. Pada tahap kedua, vektor probabilitas emosi hasil tahap pertama digabungkan (feature fusion) dengan fitur akustik dari dataset ChaLearn First Impression untuk memprediksi lima dimensi Big Five Personality (OCEAN) menggunakan Concordance Correlation Coefficient (CCC) Loss. Model estimasi kepribadian mencapai rata-rata 1-MAE sebesar 0,8964, MAE sebesar 0,1036 dan RMSE sebesar 0,1296 pada seluruh dimensi OCEAN. Hasil ablation study menunjukkan bahwa penambahan fitur emosi secara konsisten meningkatkan kinerja model pada seluruh dimensi, dengan peningkatan terbesar pada dimensi Neuroticism (+0,0106). Meskipun korelasi Pearson antara vektor emosi dan skor OCEAN tergolong rendah (|r| < 0,19), kontribusi kolektif kedelapan emosi tetap memberikan informasi yang bermakna. Hasil ini menunjukkan bahwa deteksi emosi dapat dimanfaatkan sebagai fitur antara yang bermanfaat untuk meningkatkan estimasi kepribadian berbasis suara, meskipun kinerjanya masih berada di bawah pendekatan multimodal pada ChaLearn First Impression V2 Challenge.
=================================================================================================================================
The human voice carries rich nonverbal information, including emotion and personality. Conventional personality assessment through questionnaires is subjective, motivating the need for an automatic and objective voice-based approach. This research develops a two-stage deep learning approach to estimate personality from voice data, leveraging emotion detection as an intermediate feature. However, detecting emotion from speech signals poses a particular challenge, as emotional patterns are shaped not only by local spectral characteristics such as frequency, harmonics, and intensity, but also by the temporal dynamics of the signal throughout the utterance, so a model that captures only one of these aspects tends to be suboptimal. To address this challenge, the first stage employs a CNN-BiLSTM model, where the CNN extracts local spectral patterns and the BiLSTM models bidirectional temporal dependencies, to detect eight emotion classes from mel-spectrograms using the RAVDESS dataset, achieving an accuracy of 79.17% on the testing set. In the second stage, the resulting emotion probability vector is combined through feature fusion with acoustic features from the ChaLearn First Impression dataset to predict Big Five Personality (OCEAN) dimensions using Concordance Correlation Coefficient (CCC) Loss. The personality estimation model achieves an average 1-MAE of 0.8964, MAE of 0.1036 and RMSE of 0.1296 across all OCEAN dimensions. An ablation study shows that adding emotion features consistently improves model performance across all dimensions, with the largest improvement observed for Neuroticism (+0.0106). Although Pearson correlations between the emotion vector and OCEAN scores are weak (|r| < 0.19), the collective contribution of the eight emotions still provides meaningful information. These results indicate that emotion detection can be leveraged as a useful intermediate feature to improve voice-based personality estimation, although its performance still falls below multimodal approaches in the ChaLearn First Impression V2 Challenge.

Item Type: Thesis (Other)
Uncontrolled Keywords: Big Five Personality (OCEAN), CNN-BiLSTM, Deep Learning, Estimasi Kepribadian, Speech Emotion Recognition,Big Five Personality (OCEAN), CNN-BiLSTM, Deep Learning, Personality Estimation, Speech Emotion Recognition
Subjects: Q Science > QA Mathematics > QA336 Artificial Intelligence
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Fathurazka Gamma Syuhada
Date Deposited: 31 Jul 2026 02:15
Last Modified: 31 Jul 2026 02:15
URI: http://repository.its.ac.id/id/eprint/140436

Actions (login required)

View Item View Item