Analisis Komparatif Arsitektur CNN dan Transformer untuk Deteksi Deepfake Biner pada Citra Wajah

Daniswara, M Fadhil Abhista (2026) Analisis Komparatif Arsitektur CNN dan Transformer untuk Deteksi Deepfake Biner pada Citra Wajah. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221208_undergraduate_thesis.pdf] Text
5025221208_undergraduate_thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (9MB) | Request a copy

Abstract

Perkembangan pesat teknologi Generative AI telah memicu peningkatan penyebaran citra manipulasi wajah atau *deepfake*, yang menimbulkan ancaman serius terhadap integritas informasi digital. Metode deteksi konvensional yang mengandalkan arsitektur Convolutional Neural Network (CNN) sering kali rentan mengalami *overfitting* pada generator spesifik dan kesulitan menangkap inkonsistensi spasial secara global. Penelitian ini bertujuan melakukan studi komparatif secara komprehensif antara arsitektur CNN, yang direpresentasikan oleh Residual Network-50 (ResNet-50), dan arsitektur berbasis *attention*, yaitu Vision Transformer (ViT-Base/16), dalam studi kasus klasifikasi biner citra *deepfake*. Metodologi penelitian dirancang melalui lima skenario pengujian bertahap, mulai dari pelatihan batas dasar (*baseline*), optimasi parameter (*tuning*), penskalaan data awal, integrasi fungsi objektif *Focal Loss*, hingga uji batas skala maksimal. Kinerja komprehensif model dievaluasi menggunakan metrik akurasi, presisi, *recall*, dan F1-score. Hasil eksperimen membuktikan bahwa arsitektur ViT-Base/16 secara konsisten mengungguli ResNet-50 pada seluruh metrik evaluasi. Kinerja paling optimal dicapai pada Skenario 4 (integrasi *Focal Loss* dan skala data optimal), dengan ViT-Base/16 meraih akurasi tertinggi sebesar 83,86%, dibandingkan 80,44% pada ResNet-50. Analisis lebih lanjut menunjukkan bahwa ResNet-50 memiliki sensitivitas yang baik terhadap kelas *Fake*, namun mencatat tingkat *false positive* yang tinggi karena sering keliru memprediksi kelas *Real*. Di sisi lain, penambahan kuantitas data hingga batas maksimal (Skenario 5) justru memicu *noise* yang mendegradasi performa kedua model. Meskipun ViT-Base/16 terbukti lebih unggul dalam representasi fitur global dan kemampuan generalisasi, evaluasi komputasi menunjukkan bahwa ResNet-50 memiliki efisiensi waktu inferensi yang lebih cepat, yaitu 6,14 ms per gambar, dibandingkan ViT-Base/16 yang memerlukan 10,28 ms per gambar. Penelitian ini menyimpulkan bahwa Vision Transformer sangat direkomendasikan untuk sistem deteksi *deepfake* yang menuntut keandalan akurasi tinggi, sementara CNN tetap relevan untuk implementasi sistem *real-time* dengan keterbatasan sumber daya komputasi.
=================================================================================================================================
The rapid development of Generative AI technology has triggered a significant increase in the spread of facial manipulation images, commonly known as *deepfakes*, posing a serious threat to the integrity of digital information. Conventional detection methods relying on Convolutional Neural Network (CNN) architectures are often susceptible to *overfitting* on specific generators and struggle to capture global spatial inconsistencies. This study aims to conduct a comprehensive comparative analysis between a CNN architecture, represented by Residual Network-50 (ResNet-50), and an attention-based architecture, Vision Transformer (ViT-Base/16), for the binary classification of *deepfake* images. The research methodology was designed through five progressive testing scenarios, beginning with baseline training, followed by parameter optimization (*tuning*), initial data scaling, integration of the *Focal Loss* objective function, and finally maximum-scale testing. Model performance was comprehensively evaluated using Accuracy, Precision, Recall, and F1-score metrics. The experimental results demonstrated that the ViT-Base/16 architecture consistently outperformed ResNet-50 across all evaluation metrics. The best performance was achieved in Scenario 4, which combined *Focal Loss* integration with the optimal data scale, where ViT-Base/16 attained the highest accuracy of 83.86%, compared with 80.44% for ResNet-50. Further analysis revealed that ResNet-50 exhibited good sensitivity in detecting the *Fake* class but produced a relatively high *false positive* rate due to frequent misclassification of the *Real* class. Conversely, increasing the dataset to its maximum scale in Scenario 5 introduced additional *noise*, leading to performance degradation in both models. Although ViT-Base/16 demonstrated superior capability in global feature representation and generalization, computational evaluation indicated that ResNet-50 achieved faster inference efficiency, requiring only 6.14 ms per image compared with 10.28 ms per image for ViT-Base/16. This study concludes that Vision Transformer is highly recommended for *deepfake* detection systems requiring high detection accuracy and robustness, whereas CNN-based models remain a practical choice for real-time applications with limited computational resources.

Item Type: Thesis (Other)
Uncontrolled Keywords: Deteksi Deepfake, Convolutional Neural Network, ResNet, Vision Transformer, Klasifikasi Biner. Deepfake Detection, Convolutional Neural Network, ResNet, Vision Transformer, Binary Classification
Subjects: T Technology > T Technology (General)
T Technology > T Technology (General) > T57.5 Data Processing
T Technology > TA Engineering (General). Civil engineering (General) > TA1637 Image processing--Digital techniques. Image analysis--Data processing.
T Technology > TA Engineering (General). Civil engineering (General) > TA1650 Face recognition. Optical pattern recognition.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: M Fadhil Abhista Daniswara
Date Deposited: 05 Aug 2026 05:28
Last Modified: 05 Aug 2026 05:28
URI: http://repository.its.ac.id/id/eprint/143885

Actions (login required)

View Item View Item