Priambodo, Tegar Ganang Satrio (2026) Klasifikasi Tingkat Kompleksitas Pertanyaan Visual Medis Menggunakan Fusi Representasi Teks dan Citra. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6025241037-Master_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (32MB) | Request a copy |
Abstract
Perkembangan sistem Visual Question Answering (VQA) medis membuka peluang baru dalam pemanfaatan kecerdasan buatan untuk mendukung pengambilan keputusan klinis melalui pemahaman multimodal antara citra medis dan pertanyaan bahasa alami. Meskipun demikian, mayoritas penelitian VQA saat ini masih berfokus secara eksklusif pada peningkatkan akurasi jawaban, sementara aspek yang menentukan tingkat kedalaman penalaran klinis belum banyak dieksplorasi. Pemahaman terhadap struktur pertanyaan ini sangat penting untuk membangun sistem VQA yang adaptif dan kontekstual. Tantangan utamanya adalah bagaimana sistem dapat mengklasifikasikan pertanyaan tidak hanya berdasarkan tingkat kompleksitas, tetapi juga mengevaluasinya berdasarkan level kognitif melalui Taksonomi Bloom serta kategori kelas pertanyaan spesifik terkait domain medis. Untuk menjawab tantangan tersebut, penelitian ini mengusulkan sebuah kerangka kerja pembelajaran multimodal yang memfusikan representasi teks dan citra endoskopi guna mengklasifikasikan tingkat kompleksitas (level 1, 2, dan 3), level kognitif (Remember, Analyze, Apply dan Evaluate), dan 17 kelas pertanyaan (Contoh: Abnormality Color, Polyp Count dll) visual medis. Penelitian ini memanfaatkan dataset Kvasir-VQA-x1 berskala besar yang mencakup 6.500 citra endoskopi dan 159.549 pasangan pertanyaan-jawaban yang telah dikurasi. Proses ekstraksi fitur dilakukan menggunakan kombinasi arsitektur pra-latih Convolutional Neural Network (CNN) untuk menangkap fitur visual dan representasi Transformer (seperti BERT) untuk mengekstraksi informasi semantik tekstual. Vektor dari kedua modalitas tersebut kemudian digabungkan dan dievaluasi melalui tiga strategi fusi utama Concatenation Fusion, Adaptive Attention Fusion, dan Cross Attention Fusion sebelum diproses oleh algoritma klasifikasi Machine Learning. Hasil eksperimen menunjukkan bahwa metode Cross Attention Fusion memberikan kinerja klasifikasi yang paling unggul dan konsisten dibandingkan pendekatan fusi lainnya. Kombinasi model MobileNetV2, BERT, dan Support Vector Machine (SVM) mampu mencapai tingkat akurasi hingga 98,89% pada pengujian klasifikasi tingkat kompleksitas. Selain itu, model ini juga menunjukkan hasil yang sangat kompetitif dalam memetakan level kognitif dan kelas pertanyaan, di mana evaluasi membuktikan bahwa modalitas teks memiliki peran yang jauh lebih dominan dalam menangkap beban kognitif dibandingkan modalitas visual secara mandiri. Didukung oleh analisis kualitatif menggunakan attention heatmap, temuan penelitian ini menegaskan bahwa pendekatan fusi multimodal saling melengkapi dan sangat potensial dalam mendukung pengembangan sistem VQA medis yang presisi, tangguh, dan dapat diinterpretasikan secara klinis.
================================================================================================================================
Recent developments in Visual Question Answering (VQA) systems within the field of medicine are opening up new opportunities for the use of artificial intelligence to support clinical decision-making through multimodal understanding of medical images and natural language queries. However, most current VQA research still focuses exclusively on improving answer accuracy, whilst aspects that determine the depth of clinical reasoning have not yet been extensively explored. An understanding of this question structure is crucial for building adaptive and context-aware VQA systems. The main challenge is how a system can automatically classify questions not only based on their level of complexity, but also evaluate them according to cognitive levels using Bloom’s Taxonomy, as well as specific question categories relevant to the medical domain. To address these challenges, this study proposes a multimodal learning framework that fuses text representations and endoscopic images to classify medical visuals by complexity levels (levels 1, 2, and 3), cognitive levels (Remember, Analyze, Apply, and Evaluate), and 17 question classes (Examples: Abnormality Color, Polyp Count, etc.). This study utilizes the large Kvasir-VQA-x1 dataset, which includes 6,500 endoscopic images and 159,549 curated question-answer pairs. The feature extraction process is performed using a combination of pre-trained Convolutional Neural Network (CNN) architectures to capture visual features and Transformer representations (such as BERT) to extract textual semantic information. Vectors from both modalities are then combined and evaluated through three main fusion strategies: Concatenation Fusion, Adaptive Attention Fusion, and Cross Attention Fusion before being processed by a Machine Learning classification algorithm.
Experimental results show that Cross Attention Fusion method delivers most consistent classification performance compared to other fusion approaches. The combination of MobileNetV2 model, BERT and Support Vector Machine (SVM) achieved an accuracy rate of up to 98.89% in complexity-level classification test. Furthermore, this model also demonstrated highly competitive results in mapping cognitive levels and question classes, with evaluations proving that text modality plays a far more dominant role in capturing cognitive load than visual modality alone. Supported by qualitative analysis using attention heatmaps, findings of this study confirm that multimodal fusion approach is complementary and holds great potential for supporting the development of precise, robust and clinically interpretable medical VQA systems.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | Citra Endoskopi, Pembelajaran Multimodal, Klasifikasi Kompleksitas Pertanyaan, Representasi Fitur Teks dan Citra, Visual Question Answering Medis, Endoscopic Images, Medical Visual Question Answering, Multimodal Learning, Question Complexity Classification, Text and Image Representation |
| Subjects: | Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. T Technology > T Technology (General) T Technology > T Technology (General) > T57.5 Data Processing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis |
| Depositing User: | Tegar Ganang Satrio Priambodo |
| Date Deposited: | 30 Jul 2026 01:10 |
| Last Modified: | 30 Jul 2026 01:10 |
| URI: | http://repository.its.ac.id/id/eprint/139704 |
Actions (login required)
![]() |
View Item |
