Kendali Kursi Roda Berbasis Eye Gesture Menggunakan Model Hybrid EAR–LSTM Pada Platform Komputasi Edge

Rijal, Ahmad Akmal (2026) Kendali Kursi Roda Berbasis Eye Gesture Menggunakan Model Hybrid EAR–LSTM Pada Platform Komputasi Edge. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5024221017-Undergraduate_Thesis.pdf] Text
5024221017-Undergraduate_Thesis.pdf
Restricted to Repository staff only

Download (17MB) | Request a copy

Abstract

Penyandang disabilitas motorik berat, seperti penderita tetraplegia dan *Amyotrophic Lateral Sclerosis* (ALS), umumnya mengalami keterbatasan dalam mengoperasikan kursi roda elektrik konvensional berbasis *joystick*. Kondisi tersebut mendorong pengembangan antarmuka kendali alternatif yang dapat dioperasikan tanpa kontak fisik. Penelitian ini mengembangkan *Eye-Controlled Wheelchair System* (ECWS), yaitu sistem kendali kursi roda elektrik berbasis gestur mata yang memungkinkan pengguna memberikan perintah navigasi melalui pola kedipan dan arah lirikan mata. Sistem mengekstraksi 12 fitur spasial-temporal pada setiap *frame* secara *real-time* menggunakan MediaPipe FaceMesh, yang meliputi *Eye Aspect Ratio* (EAR) kiri dan kanan, arah pandang tiga dimensi (*gaze yaw* dan *gaze pitch* 3D), orientasi kepala (*pitch*, *yaw*, dan *roll*), diameter iris kiri dan kanan, serta tiga fitur turunan kecepatan. Seluruh fitur disusun menjadi sekuens sepanjang 60 *frame* (2,0 detik pada 30 FPS) dan diklasifikasikan ke dalam delapan kelas gestur perintah. Tiga arsitektur *deep learning* dievaluasi secara komparatif, yaitu CNN-LSTM-Attention sebagai model utama, serta TCN dan InceptionTime sebagai model pembanding. Model dalam format ONNX dijalankan secara lokal pada NVIDIA Jetson Nano 4 GB menggunakan ONNX Runtime dengan akselerasi CUDA. Hasil klasifikasi dikirimkan melalui protokol UART dengan verifikasi *checksum* CRC8 ke mikrokontroler ESP32 untuk mengendalikan motor DC kursi roda melalui *driver* H-Bridge BTS7960B. Hasil evaluasi *offline* menunjukkan bahwa InceptionTime memperoleh akurasi tertinggi sebesar 98,32% (F1-Score Macro 98,48%), melampaui CNN-LSTM-Attention (97,02%; F1-Score Macro 97,30%) dan TCN (96,38%; F1-Score Macro 96,54%). Meskipun demikian, CNN-LSTM-Attention dipilih sebagai model *deployment* karena memiliki ukuran ONNX 2,8 kali lebih kecil (0,47 MB) dengan latensi inferensi sekitar 5,6 ms, sehingga lebih efisien untuk komputasi berkelanjutan pada perangkat *edge*. Evaluasi generalisasi lintas subjek tanpa kalibrasi ulang menghasilkan rata-rata akurasi sebesar 88,33% untuk model CNN-LSTM-Attention (S01: 94,38%; S06: 86,25%; S09: 84,38%). Pada pengujian terpadu secara *real-time*, sistem mencapai *command accuracy* sebesar 91,3% dari 160 percobaan pada delapan kelas gestur. Latensi komputasi *pipeline* inferensi tercatat sebesar 59,67 ms, sedangkan total latensi *end-to-end* berkisar antara 107,29 ms hingga 143,00 ms bergantung pada konfigurasi kamera, dengan proses akuisisi *frame* sebagai hambatan utama. Mekanisme keselamatan berlapis yang terdiri atas EAR Hard Override (<50 ms), *majority voting* tiga *frame*, *cooldown* selama 55 *frame*, validasi paket CRC8, dan *hardware watchdog* 300 ms berhasil menurunkan *False Command Rate* menjadi 3,0%.
=================================================================================================================================
Individuals with severe motor disabilities, such as those with tetraplegia and *Amyotrophic Lateral Sclerosis* (ALS), are generally unable to operate conventional joystick-based electric wheelchairs. This limitation necessitates the development of alternative control interfaces that can be operated without physical contact. This study presents an *Eye-Controlled Wheelchair System* (ECWS), an eye gesture-based control system that enables users to issue wheelchair navigation commands through intentional eye blink patterns and gaze direction. The system extracts 12 spatial-temporal features from each video frame in real time using MediaPipe FaceMesh, including the left and right *Eye Aspect Ratio* (EAR), three-dimensional gaze direction (*gaze yaw* and *gaze pitch*), head orientation (*pitch*, *yaw*, and *roll*), left and right iris diameters, and three velocity derivative features. These features are organized into a 60-frame sliding window (2.0 seconds at 30 FPS) and classified into eight gesture command classes. Three deep learning architectures are comparatively evaluated: CNN-LSTM-Attention as the primary model, with TCN and InceptionTime serving as baseline models. The ONNX models are executed locally on an NVIDIA Jetson Nano 4 GB using ONNX Runtime with CUDA acceleration. Classification outputs are transmitted via the UART protocol with CRC8 checksum verification to an ESP32 microcontroller, which controls the wheelchair DC motors through a BTS7960B H-Bridge motor driver. Offline evaluation shows that InceptionTime achieves the highest accuracy of 98.32% (Macro F1-Score of 98.48%), outperforming CNN-LSTM-Attention (97.02%; Macro F1-Score 97.30%) and TCN (96.38%; Macro F1-Score 96.54%). Nevertheless, CNN-LSTM-Attention is selected as the deployment model because its ONNX model size is 2.8 times smaller (0.47 MB) and its inference latency is approximately 5.6 ms, ensuring efficient long-term computation on an edge device. Cross-subject evaluation without recalibration yields an average accuracy of 88.33% for the deployed CNN-LSTM-Attention model (S01: 94.38%; S06: 86.25%; S09: 84.38%). In integrated real-time testing, the system achieves a command accuracy of 91.3% over 160 trials across eight gesture classes. The inference pipeline requires 59.67 ms of computation time, while the total end-to-end latency ranges from 107.29 ms to 143.00 ms depending on camera configuration, with frame acquisition identified as the primary bottleneck. A multi-layer safety mechanism consisting of EAR Hard Override (<50 ms), three-frame majority voting, a 55-frame cooldown period, CRC8 packet validation, and a 300 ms hardware watchdog successfully reduces the False Command Rate to 3.0%.

Item Type: Thesis (Other)
Uncontrolled Keywords: Eye Gesture Recognition, Hybrid EAR–LSTM, Edge AI, MediaPipe FaceMesh,Kursi Roda Elektrik, Safety Filter, Teknologi Asistif ============================================================= Eye Gesture Recognition, Hybrid EAR–LSTM, Edge AI, MediaPipe FaceMesh, Electric Wheelchair, Safety Filter, Assistive Technology
Subjects: T Technology > TA Engineering (General). Civil engineering (General) > TA1637 Image processing--Digital techniques. Image analysis--Data processing.
T Technology > TA Engineering (General). Civil engineering (General) > TA1650 Face recognition. Optical pattern recognition.
T Technology > TJ Mechanical engineering and machinery > TJ211.4 Robot motion
T Technology > TJ Mechanical engineering and machinery > TJ211.415 Mobile robots
T Technology > TJ Mechanical engineering and machinery > TJ212 Control engineering systems. Automatic machinery (General)
T Technology > TJ Mechanical engineering and machinery > TJ213 Automatic control.
T Technology > TJ Mechanical engineering and machinery > TJ223.A25 Actuators.
T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK7882.P3 Pattern recognition systems
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Computer Engineering > 90243-(S1) Undergraduate Thesis
Depositing User: Ahmad Akmal Rijal
Date Deposited: 22 Jul 2026 05:33
Last Modified: 22 Jul 2026 05:33
URI: http://repository.its.ac.id/id/eprint/135697

Actions (login required)

View Item View Item