Pengembangan Metode Deteksi Kecelakaan Lalu Lintas Berbasis Video Menggunakan Vision-Language Models (VLM)

Ramadhani, Syahmirza Ahmad (2026) Pengembangan Metode Deteksi Kecelakaan Lalu Lintas Berbasis Video Menggunakan Vision-Language Models (VLM). Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5024221071-Undergraduate_Thesis.pdf] Text
5024221071-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (32MB) | Request a copy

Abstract

Potensi yang besar ditunjukkan oleh Vision-Language Models (VLM) dalam tugas analisis visual yang kompleks, akan tetapi penelitian yang lebih mendalam dibutuhkan untuk diterapkan pada deteksi kejadian dinamis terutama pada insiden kritis seperti kecelakaan lalu lintas dalam format video. Oleh karena hal tersebut, tujuan dari penelitian kali ini yaitu untuk menganalisis secara komprehensif kapabilitas VLM modern pada domain deteksi kejadian dengan studi kasus utama pada kecelakaan lalu lintas menggunakan model Qwen2.5-VL dan Qwen3-VL. Metodologi penelitian kali ini mencakup empat tahapan utama yang diawali dengan menguji model terhadap dataset video kejadian guna membangun benchmark kinerja dasar (zero-shot) VLM, kemudian tahap berikutnya yaitu pengujian variasi strategi prompting untuk menganalisis sensitivitas model terhadap prompt yang diberikan , selanjutnya diikuti dengan implementasi dan evaluasi untuk mengadaptasi few-shot learning dengan metode fine-tuning efisien parameter menggunakan LoRA pada model, dan terakhir penelitian ditutup dengan menganalisis pertukaran (trade-off) antara semua pendekatan dan metode yang digunakan selama penelitian. Diharapkan hasil dari penelitian kali ini dapat memberikan data mengenai kemampuan VLM dalam mendeteksi kejadian dalam format video, Memahami variasi strategi prompting paling efektif pada model, Memberikan bukti kuantitatif trade-off dari model yang menggunakan metode zero-shot dan few-shot, dan Kesimpulan serta saran untuk penelitian atau penerapan VLM dimasa depan pada tugas serupa.
======================================================================================================================================
Vision-Language Models (VLM) have shown great potential in complex visual analysis tasks, However, more in-depth research is needed to apply them to dynamic event detection, especially in critical incidents such as traffic accidents in video format. Therefore, the purpose of this study is to comprehensively analyze the capabilities of modern VLMs in the event detection domain with a main case study on traffic accidents using the Qwen2.5-VL and Qwen3-VL model. The current research methodology includes four main stages, starting with testing the model against an incident video dataset to establish a baseline performance benchmark (zero-shot) for VLM, followed by testing variations of the prompting strategy to analyze the model's sensitivity to the prompt given by, followed by implementation and evaluation to adapt few-shot learning with a parameter-efficient fine-tuning method using LoRA on the model, and finally, the study concludes by analyzing the trade-offs between all approaches and methods used during the study. It is expected that the results of this study can provide data regarding the ability of VLM in detecting events in video format, Understanding the most effective variations of prompting strategies in the model, Providing quantitative evidence of trade-off from models using zero-shot and few-shot methods, and Conclusions and suggestions for future research or application of VLM on similar tasks.

Item Type: Thesis (Other)
Uncontrolled Keywords: Vision-Language Model (VLM), Zero-Shot Learning, Few-Shot Learning, Prompt Engineering, Qwen2.5-VL, Qwen3-VL, Deteksi Kecelakaan Lalu Lintas, Traffic Accident Detection
Subjects: H Social Sciences > HE Transportation and Communications > HE5614.3.N5 Traffic accidents
Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines.
Q Science > QA Mathematics > QA336 Artificial Intelligence
Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
T Technology > TA Engineering (General). Civil engineering (General) > TA1637 Image processing--Digital techniques. Image analysis--Data processing.
T Technology > TE Highway engineering. Roads and pavements > TE228.3 Intelligent transportation systems.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Computer Engineering > 90243-(S1) Undergraduate Thesis
Depositing User: Syahmirza Ahmad Ramadhani
Date Deposited: 24 Jul 2026 14:19
Last Modified: 24 Jul 2026 14:19
URI: http://repository.its.ac.id/id/eprint/137483

Actions (login required)

View Item View Item