Rijal, Ahmad Akmal (2026) Kendali Kursi Roda Berbasis Eye Gesture Menggunakan Model Hybrid EAR–LSTM Pada Platform Komputasi Edge. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5024221017-Undergraduate_Thesis.pdf Restricted to Repository staff only Download (17MB) | Request a copy |
Abstract
Penyandang disabilitas motorik berat, seperti penderita tetraplegia dan *Amyotrophic Lateral Sclerosis* (ALS), umumnya mengalami keterbatasan dalam mengoperasikan kursi roda elektrik konvensional berbasis *joystick*. Kondisi tersebut mendorong pengembangan antarmuka kendali alternatif yang dapat dioperasikan tanpa kontak fisik. Penelitian ini mengembangkan *Eye-Controlled Wheelchair System* (ECWS), yaitu sistem kendali kursi roda elektrik berbasis gestur mata yang memungkinkan pengguna memberikan perintah navigasi melalui pola kedipan dan arah lirikan mata. Sistem mengekstraksi 12 fitur spasial-temporal pada setiap *frame* secara *real-time* menggunakan MediaPipe FaceMesh, yang meliputi *Eye Aspect Ratio* (EAR) kiri dan kanan, arah pandang tiga dimensi (*gaze yaw* dan *gaze pitch* 3D), orientasi kepala (*pitch*, *yaw*, dan *roll*), diameter iris kiri dan kanan, serta tiga fitur turunan kecepatan. Seluruh fitur disusun menjadi sekuens sepanjang 60 *frame* (2,0 detik pada 30 FPS) dan diklasifikasikan ke dalam delapan kelas gestur perintah. Tiga arsitektur *deep learning* dievaluasi secara komparatif, yaitu CNN-LSTM-Attention sebagai model utama, serta TCN dan InceptionTime sebagai model pembanding. Model dalam format ONNX dijalankan secara lokal pada NVIDIA Jetson Nano 4 GB menggunakan ONNX Runtime dengan akselerasi CUDA. Hasil klasifikasi dikirimkan melalui protokol UART dengan verifikasi *checksum* CRC8 ke mikrokontroler ESP32 untuk mengendalikan motor DC kursi roda melalui *driver* H-Bridge BTS7960B. Hasil evaluasi *offline* menunjukkan bahwa InceptionTime memperoleh akurasi tertinggi sebesar 98,32% (F1-Score Macro 98,48%), melampaui CNN-LSTM-Attention (97,02%; F1-Score Macro 97,30%) dan TCN (96,38%; F1-Score Macro 96,54%). Meskipun demikian, CNN-LSTM-Attention dipilih sebagai model *deployment* karena memiliki ukuran ONNX 2,8 kali lebih kecil (0,47 MB) dengan latensi inferensi sekitar 5,6 ms, sehingga lebih efisien untuk komputasi berkelanjutan pada perangkat *edge*. Evaluasi generalisasi lintas subjek tanpa kalibrasi ulang menghasilkan rata-rata akurasi sebesar 88,33% untuk model CNN-LSTM-Attention (S01: 94,38%; S06: 86,25%; S09: 84,38%). Pada pengujian terpadu secara *real-time*, sistem mencapai *command accuracy* sebesar 91,3% dari 160 percobaan pada delapan kelas gestur. Latensi komputasi *pipeline* inferensi tercatat sebesar 59,67 ms, sedangkan total latensi *end-to-end* berkisar antara 107,29 ms hingga 143,00 ms bergantung pada konfigurasi kamera, dengan proses akuisisi *frame* sebagai hambatan utama. Mekanisme keselamatan berlapis yang terdiri atas EAR Hard Override (<50 ms), *majority voting* tiga *frame*, *cooldown* selama 55 *frame*, validasi paket CRC8, dan *hardware watchdog* 300 ms berhasil menurunkan *False Command Rate* menjadi 3,0%.
=================================================================================================================================
Individuals with severe motor disabilities, such as those with tetraplegia and *Amyotrophic Lateral Sclerosis* (ALS), are generally unable to operate conventional joystick-based electric wheelchairs. This limitation necessitates the development of alternative control interfaces that can be operated without physical contact. This study presents an *Eye-Controlled Wheelchair System* (ECWS), an eye gesture-based control system that enables users to issue wheelchair navigation commands through intentional eye blink patterns and gaze direction. The system extracts 12 spatial-temporal features from each video frame in real time using MediaPipe FaceMesh, including the left and right *Eye Aspect Ratio* (EAR), three-dimensional gaze direction (*gaze yaw* and *gaze pitch*), head orientation (*pitch*, *yaw*, and *roll*), left and right iris diameters, and three velocity derivative features. These features are organized into a 60-frame sliding window (2.0 seconds at 30 FPS) and classified into eight gesture command classes. Three deep learning architectures are comparatively evaluated: CNN-LSTM-Attention as the primary model, with TCN and InceptionTime serving as baseline models. The ONNX models are executed locally on an NVIDIA Jetson Nano 4 GB using ONNX Runtime with CUDA acceleration. Classification outputs are transmitted via the UART protocol with CRC8 checksum verification to an ESP32 microcontroller, which controls the wheelchair DC motors through a BTS7960B H-Bridge motor driver. Offline evaluation shows that InceptionTime achieves the highest accuracy of 98.32% (Macro F1-Score of 98.48%), outperforming CNN-LSTM-Attention (97.02%; Macro F1-Score 97.30%) and TCN (96.38%; Macro F1-Score 96.54%). Nevertheless, CNN-LSTM-Attention is selected as the deployment model because its ONNX model size is 2.8 times smaller (0.47 MB) and its inference latency is approximately 5.6 ms, ensuring efficient long-term computation on an edge device. Cross-subject evaluation without recalibration yields an average accuracy of 88.33% for the deployed CNN-LSTM-Attention model (S01: 94.38%; S06: 86.25%; S09: 84.38%). In integrated real-time testing, the system achieves a command accuracy of 91.3% over 160 trials across eight gesture classes. The inference pipeline requires 59.67 ms of computation time, while the total end-to-end latency ranges from 107.29 ms to 143.00 ms depending on camera configuration, with frame acquisition identified as the primary bottleneck. A multi-layer safety mechanism consisting of EAR Hard Override (<50 ms), three-frame majority voting, a 55-frame cooldown period, CRC8 packet validation, and a 300 ms hardware watchdog successfully reduces the False Command Rate to 3.0%.
Actions (login required)
![]() |
View Item |
