Abdillah, M. Andi (2026) Pengembangan Sistem Instruksi Suara Adaptif Berbasis LLM-Vision Reasoning Untuk Membantu Tunanetra Mengoperasikan Remote AC Yang Belum Dikenal. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5022221073-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (4MB) | Request a copy |
Abstract
Pengoperasian remote AC secara mandiri masih menjadi tantangan bagi pengguna tunanetra karena tombol tidak memiliki penanda taktil dan tata letaknya bervariasi antarmerek. Penelitian ini mengembangkan sistem instruksi suara adaptif berbasis Vision-Language Model (VLM) untuk memandu pergerakan ibu jari menuju tombol target pada remote AC yang belum dikenal. Sistem menggunakan arsitektur client-server dengan WebRTC untuk transmisi video dan audio real-time dari smartphone ke server PC lokal, serta WebSocket untuk sinyal kontrol dan distribusi TTS. Pipeline AI menggabungkan Faster Whisper untuk transkripsi, YOLO26n OBB untuk deteksi remote dan ibu jari, OWLv2 untuk ekstraksi tata letak tombol secara zero-shot, serta Qwen3.5 9B terkuantisasi INT4 untuk pemetaan fungsi tombol dan penalaran navigasi. Hasil pengujian menunjukkan audio packet loss sebesar 0,09%, latensi sinyal kontrol 5,87 ms per perintah, dan waktu pembentukan koneksi awal 5,28 detik. YOLO26n OBB mencapai mAP@0.5 sebesar 0,995. OWLv2 menghasilkan 722 kandidat tombol dari 50 remote dengan waktu inferensi rata-rata 3,02 detik. Qwen3.5 9B memberikan label selain tidak diketahui pada 96,54% kandidat (bukan akurasi semantik) dan mencapai akurasi tugas operasional 93,2% dengan kebutuhan VRAM 6,7 GB. Faster Whisper memperoleh Word Error Rate (WER) 2,33% pada kondisi sunyi dan 13,44% pada kondisi bising dengan waktu inferensi 485 ms. Pencocokan rule-based dieksekusi dalam waktu kurang dari 1 ms, sedangkan pipeline visual VLM membutuhkan rata-rata 12,37 detik; kedua nilai tersebut memiliki batas pengukuran yang berbeda. Evaluasi formatif terhadap delapan partisipan dengan penutup mata menghasilkan Task Completion Rate 81,1% dan Task Completion Time rata-rata 15,9 detik. Pada pengujian eksploratif, tiga partisipan tunanetra akhirnya menyelesaikan 12 tugas yang dicoba dengan kamera statis, tetapi masih memerlukan percobaan ulang dan bantuan peneliti; sedikitnya dua pertanyaan tambahan mengenai hubungan spasial antartombol belum dijawab dengan tepat. Hasil ini menunjukkan kelayakan awal sistem, sedangkan kesimpulan mengenai pengalaman pengguna masih dibatasi oleh ukuran sampel dan jumlah percobaan yang kecil serta ketergantungan pada bantuan peneliti.
=================================================================================================================================
Operating AC remote controls independently remains challenging for people with visual impairments because the buttons lack tactile markers and their layouts vary across brands. This research develops an adaptive voice instruction system based on a Vision-Language Model (VLM) to guide the user’s thumb toward target buttons on unfamiliar AC remotes. The system uses a client-server architecture with WebRTC for real-time smartphone video and audio transmission to a local server, and WebSocket for control signals and TTS distribution. The AI pipeline combines Faster Whisper for transcription, YOLO26n OBB for remote and thumb detection, OWLv2 for zero-shot button layout extraction, and an INT4-quantized Qwen3.5 9B VLM for button function mapping and navigation reasoning. Experimental results show 0.09% audio packet loss, 5.87 ms control-signal latency per command, and 5.28 seconds for initial connection setup. YOLO26n OBB attains mAP@0.5 of 0.995. OWLv2 produces 722 button candidates from 50 remotes in 3.02 seconds on average. Qwen3.5 9B assigns labels other than tidak diketahui (“unknown”) to 96.54% of the candidates (not semantic accuracy) and achieves 93.2% operational task accuracy with 6.7 GB VRAM usage. Faster Whisper yields a 2.33% Word Error Rate (WER) in quiet conditions and 13.44% in noisy conditions with 485 ms inference time. Rule-based matching executes in less than 1 ms, whereas the visual VLM pipeline takes 12.37 seconds on average; these measurements use different boundaries. In a formative evaluation with eight blindfolded participants, the system achieved an 81.1% Task Completion Rate and an average Task Completion Time of 15.9 seconds. In an exploratory evaluation, three participants with visual impairments eventually completed all 12 task trials with a static camera, but still required repeated attempts and researcher assistance; at least two additional queries about spatial relationships between buttons were not answered correctly. These findings indicate initial feasibility, while conclusions about user experience remain limited by the small sample and trial count and reliance on researcher assistance.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Sistem asistif, tunanetra, remote AC, Vision-Language Model, YOLO OBB, deteksi zero-shot, panduan suara adaptif, Assistive system, people with visual impairments, AC remote, Vision-Language Model, YOLO OBB, zero-shot detection, adaptive voice guidance. |
| Subjects: | R Medicine > R Medicine (General) > R858 Deep Learning |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Electrical Engineering > 20201-(S1) Undergraduate Thesis |
| Depositing User: | M. Andi Abdillah |
| Date Deposited: | 24 Jul 2026 01:27 |
| Last Modified: | 24 Jul 2026 01:27 |
| URI: | http://repository.its.ac.id/id/eprint/135762 |
Actions (login required)
![]() |
View Item |
