Fusi Dual-Branch Fast Fourier Transform dan Vision Transformer Untuk Penilaian Estetika Foto

Rachmadan, Reza (2026) Fusi Dual-Branch Fast Fourier Transform dan Vision Transformer Untuk Penilaian Estetika Foto. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 6025231073-Master_Thesis.pdf] Text
6025231073-Master_Thesis.pdf
Restricted to Repository staff only

Download (4MB) | Request a copy

Abstract

Dalam era digital yang semakin berkembang, menilai estetika sebuah foto telah menjadi kebutuhan yang cukup penting, tidak hanya bagi individu namun juga bagi industri kreatif, platform media sosial, dan layanan e-commerce dalam menyajikan konten visual berkualitas kepada pengguna dan komsumen. Penelitian terkini tentang Image Aesthetic Assessment (IAA) menggunakan Vision Transformer (ViT) menunjukkan kemajuan yang signifikan dalam mengembangkan model yang dapat meniru penilaian estetika oleh manusia. Meskipun demikian, terdapat celah penelitian yang cukup mendasar, dimana model IAA yang ada fokus terhadap domain spasial, sementara domain lain seperti frekuensi yang memiliki informasi teknis seperti ketajaman detail citra, distribusi derau, dan pola bokeh belum pernah dieksplorasi dalam arsitektur berbasis ViT. Penelitian ini mengusulkan arsitektur Dual-Branch Fast Fourier Transform-Vision Transformer yang mengintegrasikan representasi fitur dari domain spasial dan domain frekuensi melalui mekanisme fusi konkatenasi dan adaptive gate. Arsitektur yang diusulkan terdiri atas dua cabang berbeda, yaitu cabang RGB untuk mengekstraksi informasi spasial dan cabang FFT untuk mengekstraksi informasi frekuensi. Representasi fitur dari kedua cabang kemudian digabungkan menjadi satu representasi sebelum digunakan untuk memprediksi skor estetika. Penelitian ini menggunakan Aesthetics and Attributes Database (AADB) yang terdiri atas 10.000 citra, dengan skor estetika yang diperoleh dari rata-rata penilaian lima penilai independen sebagai ground truth. Evaluasi dilakukan menggunakan metrik Pearson Linear Correlation Coefficient (PLCC), Spearman Rank-order Correlation Coefficient (SRCC), Mean Absolute Error (MAE), Mean Squared Error (MSE), dan Root Mean Squared Error (RMSE). Hasil eksperimen menunjukkan bahwa model Dual-Branch FFT-ViT dengan fusi konkatenasi memberikan performa terbaik dengan memperoleh PLCC sebesar 0,7336, SRCC sebesar 0,7347, MSE sebesar 0,0193, MAE sebesar 0,1117, dan RMSE sebesar 0,1388, serta meningkatkan nilai PLCC sebesar 0,0124 dibandingkan model domain tunggal RGB sebagai domain tunggal terbaik. Hasil tersebut menunjukkan bahwa integrasi informasi dari domain piksel dan domain frekuensi mampu menghasilkan representasi fitur yang lebih efektif untuk meningkatkan akurasi prediksi estetika dibandingkan penggunaan single domain.
==========================================================================================================================================
In the increasingly growing digital era, assessing the aesthetics of a photo has become quite an important necessity, not only for individuals but also for the creative industry, social media platforms, and e-commerce services in presenting quality visual content to users and consumers. Recent research on Image Aesthetic Assessment (IAA) using Vision Transformers (ViT) shows significant progress in developing models that can mimic aesthetic assessment by humans. However, there is a fairly fundamental research gap, where existing IAA models focus on spatial domains, while other domains such as frequencies that have technical information such as image detail sharpness, noise distribution, and bokeh patterns have never been explored in ViT-based architectures. This study proposes a Dual-Branch Fast Fourier Transform-Vision Transformer architecture that integrates feature representations of spatial domains and frequency domains through concatenation fusion mechanisms and adaptive gates. The proposed architecture consists of two distinct branches, namely the RGB branch to extract spatial information and the FFT branch to extract frequency information. The feature representations of the two branches are then combined into a single representation before being used to predict aesthetic scores. This study uses the Aesthetics and Attributes Database (AADB) which consists of 10,000 images, with an aesthetic score obtained from an average assessment of five independent assessors as ground truth. The evaluation was conducted using the Pearson Linear Correlation Coefficient (PLCC), Spearman Rank-order Correlation Coefficient (SRCC), Mean Absolute Error (MAE), Mean Squared Error (MSE), and Root Mean Squared Error (RMSE) metrics. The results of the experiment showed that the Dual-Branch FFT-ViT model with concatenation fusion provided the best performance by obtaining PLCC of 0.7336, SRCC of 0.7347, MSE of 0.0193, MAE of 0.1117, and RMSE of 0.1388, and increased the PLCC value of 0.0124 compared to the RGB single-domain model as the best single domain. These results show that the integration of information from pixel domains and frequency domains is able to produce more effective feature representations to improve the accuracy of aesthetic predictions compared to the use of a single domain.

Item Type: Thesis (Masters)
Uncontrolled Keywords: CNN, Dataset AADB, Domain Frekuensi, Domain Spasial, Estetika Foto, FFT, Pembelajaran Mendalam, Penilaian Estetika Foto, Foto RGB, Visual Transformer, AADB Dataset, CNN, Deep Learning, Fast Fourier Transform, Frequency Domain, Photo Aesthetic Assessment, RGB Photo, Spatial Domain, Visual Transformer
Subjects: T Technology > T Technology (General) > T385 Visualization--Technique
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: Reza Rachmadan
Date Deposited: 31 Jul 2026 03:14
Last Modified: 31 Jul 2026 03:14
URI: http://repository.its.ac.id/id/eprint/140833

Actions (login required)

View Item View Item