Harvianti, Azizah Elok (2026) Klasifikasi Emosi Musik Piano Menggunakan Pendekatan Multimodal Dengan Perbandingan Teknik Fusion. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221243-Undergraduate_Thesis.pdf Restricted to Repository staff only Download (5MB) | Request a copy |
Abstract
Penelitian ini mengusulkan pendekatan multimodal late fusion untuk klasifikasi emosi musik piano pada dataset EMOPIA berdasarkan Russell Circumplex Model dengan empat kuadran: Q1 (Happy), Q2 (Angry), Q3 (Sad), Q4 (Peaceful). Pendekatan unimodal yang hanya memanfaatkan audio atau MIDI saja memiliki keterbatasan dalam menangkap nuansa emosi musik yang kompleks, sehingga penggabungan kedua modalitas berpotensi menghasilkan representasi yang lebih lengkap. Encoder audio menggunakan arsitektur Short-Chunk ResNet, sedangkan encoder MIDI menggunakan BiLSTM dengan Structured Self-Attention. Tiga varian teknik late fusion yang dibandingkan, yaitu Addition Fusion, Concatenation Fusion, dan Weighted Average Fusion, dengan strategi pretraining dan selective freezing encoder yang mempertahankan bobot encoder unimodal pada tahap fusion. Evaluasi dilakukan melalui 5-fold cross-validation dengan set uji terkunci pada 999 klip subset EMOPIA hasil validasi manual. Hasil eksperimen menunjukkan Concatenation Fusion sebagai konfigurasi terbaik dengan accuracy 0,670 dan F1-macro 0,660, mengungguli baseline unimodal Hung dkk. (2021) yang menghasillkan audio dengan F1-macro 0,613 dan F1-macro MIDI sebesar 0,508. Studi ablasi menunjukkan modalitas audio berperan dominan dengan penurunan F1-macro sebesar 0,098 saat dihilangkan, sedangkan modalitas MIDI bersifat pelengkap dengan penurunan 0,035. Concatenation Fusion juga mengungguli dua metode state-of-the-art (SOTA) ketika dievaluasi pada pipeline yang sama, yaitu Zhao & Yoshii (2023) dengan F1-macro 0,615 dan Xiao dkk. (2025) BFAM dengan F1-macro 0,581. Arsitektur late fusion sederhana dengan strategi pretraining dan selective freezing menjadi alternatif efektif dibanding arsitektur SOTA yang lebih kompleks untuk dataset berukuran sedang seperti EMOPIA.
=================================================================================================================================
This research proposes a multimodal late fusion approach for piano music emotion classification on the EMOPIA dataset based on Russell's Circumplex Model with four quadrants: Q1 (Happy), Q2 (Angry), Q3 (Sad), Q4 (Peaceful). Unimodal approaches that utilize only audio or MIDI features have limitations in capturing the complex nuances of musical emotion, making the combination of both modalities promising for producing richer representations. The audio encoder uses the Short-Chunk ResNet architecture, while the MIDI encoder uses BiLSTM with Structured Self-Attention. Three late fusion variants are compared, namely Addition Fusion, Concatenation Fusion, and Weighted Average Fusion, with a pretraining and selective freezing strategy that preserves the unimodal encoder weights during the fusion stage. Evaluation is conducted through 5-fold cross-validation with a locked test set on 999 clips from a manually validated EMOPIA subset. The experimental results show Concatenation Fusion as the best configuration with an accuracy of 0.670 and F1-macro of 0.660, outperforming the unimodal baselines of Hung et al. (2021) for audio with F1-macro of 0.613 and MIDI F1-macro 0.508. The ablation study indicates that the audio modality plays a dominant role with the F1-macro decreasing by 0.098 when removed, while the MIDI modality serves as a complement with a decrease of 0.035. Concatenation Fusion also outperforms two state-of-the-art (SOTA) methods when evaluated on the same pipeline, namely Zhao & Yoshii (2023) with F1-macro 0.615 and Xiao et al. (2025) BFAM with F1-macro 0.581. A simple late fusion architecture with a pretraining and selective freezing strategy serves as an effective alternative compared to more complex SOTA architectures for medium-sized datasets such as EMOPIA.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Audio, EMOPIA, Klasifikasi Emosi Musik, Late Fusion, MIDI, Multimodal, Russell Circumplex Model, Music Emotion Recognition |
| Subjects: | M Music and Books on Music > M Music T Technology > T Technology (General) T Technology > T Technology (General) > T57.5 Data Processing T Technology > T Technology (General) > T58.5 Information technology. IT--Auditing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Azizah Elok Harvianti |
| Date Deposited: | 24 Jul 2026 01:33 |
| Last Modified: | 24 Jul 2026 01:33 |
| URI: | http://repository.its.ac.id/id/eprint/136881 |
Actions (login required)
![]() |
View Item |
