Pengembangan Kerangka Kerja Deteksi Kebohongan Multimodal Berbasis Video dengan Multi-Scale Temporal Graph Network dan Fusi Evidensial.

Rahayu, Yeni Dwi (2026) Pengembangan Kerangka Kerja Deteksi Kebohongan Multimodal Berbasis Video dengan Multi-Scale Temporal Graph Network dan Fusi Evidensial. Doctoral thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 7025231007-Doctoral.pdf] Text
7025231007-Doctoral.pdf - Accepted Version
Restricted to Repository staff only

Download (4MB) | Request a copy

Abstract

Kemampuan manusia dalam mendeteksi kebohongan rata-rata hanya mencapai 54–56%, tidak berbeda signifikan dari tebakan acak, bahkan pada penilai terlatih. Pendekatan konvensional seperti poligraf juga menghadapi keterbatasan fundamental, terutama kerentanan terhadap countermeasures dan ketergantungan tinggi pada interpretasi operator. Penelitian ini mengembangkan kerangka kerja deteksi kebohongan multimodal berbasis video yang terstruktur, dapat direproduksi, dan dapat diaudit. Berdasarkan tinjauan sistematis terhadap 42 studi periode 2019–2024, ditemukan empat kesenjangan utama: belum tersedianya dataset deteksi kebohongan berbasis video dari Indonesia, belum adanya pemodelan temporal graph berbasis dynamic adjacency untuk deteksi kebohongan level video dari behavioral landmark, belum adanya explainability yang komprehensif pada sebagian besar model, dan terbatasnya evaluasi reliabilitas. Penelitian ini menghasilkan lima kontribusi. Pertama, tinjauan literatur sistematis yang memetakan perkembangan, kesenjangan, dan arah penelitian deteksi kebohongan berbasis video. Kedua, Indonesian Deception Detection Dataset (I3D), dataset multimodal Indonesia dengan 1.568 video dari 196 partisipan lintas enam kelompok etnis. Ketiga, MSTGNet, model multi-scale temporal graph network yang mencapai AUC 0,835 pada I3D dan 0,799 pada dataset Real-Life Trial (RLT). Keempat, HIVE-Fusion, kerangka fusi hierarkis evidensial berbasis Dempster–Shafer yang mengintegrasikan modalitas visual, audio, dan teks serta menghasilkan keluaran yang dapat diaudit, mencapai AUC 0,9141 pada I3D. Kelima, evaluasi reliabilitas dan akuntabilitas mengungkapkan miskalibrasi sedang, kesalahan berkepercayaan tinggi, dan keterbatasan transfer lintas dataset. Temuan ini menunjukkan bahwa sistem paling tepat diposisikan sebagai alat bantu analisis dalam kerangka human-in-the-loop, bukan sebagai pengambil keputusan otomatis, serta memerlukan validasi lokal sebelum diterapkan pada konteks berdampak tinggi.
==================================================================================================================================
Human deception detection accuracy averages only 54–56%, statistically indistinguishable from chance, even among trained professionals. Conventional approaches such as the polygraph are also constrained by fundamental limitations, particularly susceptibility to countermeasures and high dependence on operator interpretation. This dissertation develops a structured, reproducible, and auditable framework for video-based multimodal deception detection. A systematic review of 42 studies published between 2019 and 2024 identified four major gaps: the absence of an Indonesian video-based deception dataset, the lack of dynamic-adjacency temporal graph modeling for video-level deception detection from behavioral landmarks, the lack of comprehensive explainability in most models, and limited reliability evaluation. This dissertation makes five main contributions. First, it provides a systematic literature review that maps current progress, gaps, and directions in video-based deception detection research. Second, it introduces the Indonesian Deception Detection Dataset (I3D), a multimodal dataset containing 1,568 videos from 196 participants across six Indonesian ethnic groups. Third, it develops MSTGNet, a multi-scale temporal graph network that achieved an AUC of 0.835 on I3D and 0.799 on the Real-Life Trial (RLT) dataset. Fourth, it proposes HIVE-Fusion, a Dempster–Shafer-based evidential hierarchical fusion framework that integrates visual, audio, and textual modalities and produces auditable outputs, achieving an AUC of 0.9141 on I3D. Fifth, it evaluates the framework from a reliability and accountability perspective, revealing moderate miscalibration, high-confidence errors, and limited cross-dataset transfer. These findings indicate that the system is most appropriately positioned as an analytical support tool within a human-in-the-loop framework, rather than as an autonomous decision-maker, and requires local validation before deployment in high-stakes contexts.

Item Type: Thesis (Doctoral)
Uncontrolled Keywords: deteksi kebohongan multimodal, temporal graph network, fusi evidensial, dataset Indonesia, multimodal deception detection, temporal graph network, evidential fusion, Indonesian dataset
Subjects: T Technology > T Technology (General) > T58.62 Decision support systems
Divisions: Faculty of Electrical Technology > Electrical Engineering > 20001-(S3) PhD Thesis
Depositing User: Yeni Dwi Rahayu
Date Deposited: 01 Aug 2026 02:54
Last Modified: 01 Aug 2026 02:54
URI: http://repository.its.ac.id/id/eprint/141414

Actions (login required)

View Item View Item