Question Answering Multimodal Citra dan Audio Berbasis Model Qwen dengan Metode Retrieval-Augmented Generation

Widiprasetyo, Defender Artha (2026) Question Answering Multimodal Citra dan Audio Berbasis Model Qwen dengan Metode Retrieval-Augmented Generation. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5024221032-Undergraduate_Thesis.pdf] Text
5024221032-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (14MB) | Request a copy

Abstract

Perkembangan Large Language Model (LLM) telah mendorong pemanfaatan kecerdasan buatan dalam sistem question answering. Namun, LLM memiliki keterbatasan karena pengetahuan yang digunakan bersifat statis dan dapat menghasilkan jawaban yang tidak sesuai fakta atau hallucination. Permasalahan tersebut menjadi semakin penting pada layanan informasi yang membutuhkan jawaban berbasis dokumen lokal dan spesifik terhadap suatu institusi. Untuk mengatasi keterbatasan tersebut, metode Retrieval-Augmented Generation (RAG) digunakan dengan menggabungkan kemampuan generatif model bahasa dan mekanisme pencarian informasi dari basis pengetahuan eksternal. Penelitian ini bertujuan untuk merancang, mengimplementasikan, dan mengevaluasi sistem question answering multimodal berbasis model Qwen dengan metode RAG pada lingkungan komputasi terbatas.

Sistem yang dikembangkan menggunakan basis pengetahuan berbasis dokumen tekstual yang diproses melalui tahapan ekstraksi teks, chunking, embedding, dan pengindeksan menggunakan FAISS. Input teks diproses melalui jalur RAG, sedangkan input citra dan audio diproses menggunakan model Qwen multimodal untuk mendukung skenario pemahaman citra, pemahaman audio, serta kombinasi input multimodal dengan konteks hasil retrieval. Pengujian dilakukan pada skenario text-only RAG, image-only, audio-only, text-image RAG, dan text-audio RAG. Penelitian ini juga membandingkan metode chunking fixed-size dan recursive, serta membandingkan konfigurasi representasi numerik model berupa FP32, FP16, dan 4-bit NF4. Selain itu, dilakukan pengujian tambahan INT8 Weight-Only A16W8 pada konfigurasi fixed-size untuk menganalisis pengaruhnya terhadap latency dan kebutuhan memori GPU.

Evaluasi dilakukan menggunakan metrik BLEU-1, ROUGE-L F1, context precision, context recall, cosine similarity, latency, dan kebutuhan memori GPU. Hasil pengujian menunjukkan bahwa sistem RAG mampu mengambil konteks yang relevan dan menghasilkan jawaban yang sesuai dengan dokumen sumber pada beberapa skenario. Konfigurasi FP16 memberikan kualitas keluaran yang paling seimbang pada banyak skenario RAG, terutama pada konfigurasi fixed-size. Konfigurasi 4-bit NF4 memberikan kebutuhan memori GPU paling rendah, tetapi tidak selalu menghasilkan latency paling rendah karena adanya overhead kuantisasi dan dekuantisasi. Hasil pengujian tambahan menunjukkan bahwa INT8 Weight-Only A16W8 mampu menurunkan latency pada beberapa skenario fixed-size, meskipun penghematan memorinya masih berada di antara FP16 dan 4-bit NF4. Dengan demikian, pemilihan konfigurasi model perlu mempertimbangkan trade-off antara kualitas keluaran, waktu respons, dan kebutuhan memori GPU.
======================================================================================================================================
The development of Large Language Models (LLMs) has encouraged the use of artificial intelligence in question answering systems. However, LLMs have limitations because their knowledge is static and may produce factually incorrect answers or hallucinations. This limitation becomes more critical in information services that require answers grounded in local and institution-specific documents. To address this issue, Retrieval-Augmented Generation (RAG) is used by combining the generative capability of language models with information retrieval from an external knowledge base. This research aims to design, implement, and evaluate a multimodal question answering system based on the Qwen model and the RAG method in a limited-computing-resource environment.

The developed system uses a textual document-based knowledge base processed through text extraction, chunking, embedding, and indexing using FAISS. Text input is processed through the RAG pipeline, while image and audio inputs are processed using Qwen multimodal models to support image understanding, audio understanding, and multimodal input combined with retrieved contexts. The evaluation is conducted in several scenarios, including text-only RAG, image-only, audio-only, text-image RAG, and text-audio RAG. This research also compares fixed-size and recursive chunking methods, as well as numerical representation configurations consisting of FP32, FP16, and 4-bit NF4. In addition, an INT8 Weight-Only A16W8 evaluation is conducted on the fixed-size configuration to analyze its effect on latency and GPU memory usage.

The evaluation uses BLEU-1, ROUGE-L F1, context precision, context recall, cosine similarity, latency, and GPU memory usage metrics. The results show that the RAG system can retrieve relevant contexts and generate answers aligned with the source documents in several scenarios. FP16 provides the most balanced output quality across many RAG scenarios, especially in the fixed-size configuration. The 4-bit NF4 configuration achieves the lowest GPU memory usage, but it does not always produce the lowest latency due to quantization and dequantization overhead. The additional evaluation shows that INT8 Weight-Only A16W8 can reduce latency in several fixed-size scenarios, although its memory efficiency remains between FP16 and 4-bit NF4. Therefore, selecting the model configuration requires considering the trade-off between output quality, response time, and GPU memory usage.

Item Type: Thesis (Other)
Uncontrolled Keywords: Question Answering, Qwen, Retrieval-Augmented Generation, RAG, Multimodal, FP32, FP16, 4-bit NF4, INT8 Weight-Only. Question Answering, Qwen, Retrieval-Augmented Generation, RAG, Multimodal, FP32, FP16, 4-bit NF4, INT8 Weight-Only.
Subjects: Q Science > QA Mathematics > QA76.6 Computer programming.
Q Science > QA Mathematics > QA76.758 Software engineering
T Technology > T Technology (General) > T57.5 Data Processing
T Technology > T Technology (General) > T58.5 Information technology. IT--Auditing
Z Bibliography. Library Science. Information Resources > ZA Information resources > Z699.5 Information storage and retrieval systems
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Computer Engineering > 90243-(S1) Undergraduate Thesis
Depositing User: Defender Artha Widiprasetyo
Date Deposited: 23 Jul 2026 07:41
Last Modified: 23 Jul 2026 07:41
URI: http://repository.its.ac.id/id/eprint/136615

Actions (login required)

View Item View Item