Husnan, Febrian Abdan (2026) Pengembangan Sistem Berbasis Multimodal Deep Learning Untuk Deteksi Penyakit Mata Menggunakan Citra Fundus Retina Pada Imbalanced Dataset. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5026221117-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (7MB) | Request a copy |
Abstract
Gangguan penglihatan dan penyakit mata masih menjadi permasalahan kesehatan yang memerlukan dukungan teknologi deteksi dini yang cepat dan akurat. Salah satu pendekatan yang banyak digunakan adalah klasifikasi penyakit mata berbasis citra fundus retina menggunakan deep learning. Namun, pengembangan model klasifikasi pada data medis masih menghadapi berbagai tantangan, seperti ketidakseimbangan kelas, keterbatasan data pada kelas minoritas, kemiripan visual antarpenyakit, serta belum optimalnya pemanfaatan data nonvisual pasien. Penelitian ini bertujuan mengembangkan sistem klasifikasi delapan penyakit mata (Normal, Diabetic, Glaucoma, Cataract, Age-related Macular Degeneration, Hypertension, Pathological Myopia, dan Other Disorders) menggunakan dataset ODIR-5K dengan membandingkan model unimodal berbasis citra fundus dan model multimodal yang menggabungkan citra fundus dengan data tabular berupa usia dan jenis kelamin. Tahapan penelitian meliputi prapemrosesan data citra dan data tabular, penanganan ketidakseimbangan data menggunakan cWGAN-GP untuk menghasilkan citra sintetis dan SMOTE untuk data tabular, pelatihan beberapa arsitektur deep learning, evaluasi model, analisis confusion matrix, serta implementasi ke dalam aplikasi web berbasis Flask. Model cWGAN-GP dilatih selama 250 epoch dan menghasilkan 1.088 citra sintetis untuk memperkaya kelas minoritas. Pada eksperimen unimodal, model terbaik adalah ResNet50 dengan customized head dan optimizer RMSprop yang memperoleh akurasi sebesar 52,18%, precision-macro sebesar 46,76%, dan F1-macro sebesar 46,60%. Model multimodal dengan pendekatan late fusion meningkatkan akurasi menjadi 60,40% dan precision-macro menjadi 59,87%, tetapi F1-macro hanya mencapai 42,95%. Hasil ini menunjukkan bahwa penambahan data tabular mampu meningkatkan ketepatan prediksi secara umum, tetapi belum sepenuhnya meningkatkan kemampuan model dalam mengenali seluruh kelas secara merata, terutama kelas minoritas, karena fitur usia dan jenis kelamin belum memberikan kontribusi diagnostik yang cukup kuat serta berpotensi menimbulkan bias terhadap kelas mayoritas. Penelitian ini menghasilkan sistem klasifikasi yang membandingkan pendekatan unimodal dan multimodal, sekaligus menekankan pentingnya peningkatan kualitas data sintetis, penambahan fitur klinis yang lebih informatif, serta validasi eksternal agar model lebih siap diterapkan pada lingkungan klinis nyata.
===============================================================================================================================
Vision impairment and eye diseases remain significant health challenges that require rapid and accurate early detection technologies. One of the most widely adopted approaches is the classification of eye diseases from retinal fundus images using deep learning. However, developing classification models for medical data continues to face several challenges, including class imbalance, limited data for minority classes, visual similarity among diseases, and the suboptimal utilization of non-visual patient information. This study aims to develop a classification system for eight eye diseases (Normal, Diabetic, Glaucoma, Cataract, Age-related Macular Degeneration, Hypertension, Pathological Myopia, and Other Disorders) using the ODIR-5K dataset by comparing a unimodal model based solely on fundus images with a multimodal model that combines fundus images with tabular data consisting of age and gender. The research workflow included image and tabular data preprocessing, handling data imbalance using cWGAN-GP to generate synthetic images and SMOTE for tabular data, training several deep learning architectures, model evaluation, confusion matrix analysis, and deployment into a Flask-based web application. The cWGAN-GP model was trained for 250 epochs and generated 1,088 synthetic images to enrich the minority classes. In the unimodal experiment, the best-performing model was ResNet50 with a customized head and the RMSprop optimizer, achieving an accuracy of 52.18%, a macro-precision of 46.76%, and a macro F1-score of 46.60%. The multimodal model employing a late fusion approach improved the accuracy to 60.40% and the macro-precision to 59.87%, while the macro F1-score reached only 42.95%. These results indicate that incorporating tabular data improves overall prediction accuracy but does not fully enhance the model's ability to recognize all classes equally, particularly minority classes, because age and gender provide limited diagnostic information and may introduce bias toward the majority classes. This study presents a classification system that compares unimodal and multimodal approaches while highlighting the importance of improving the quality of synthetic data, incorporating more informative clinical features, and performing external validation to enhance the model's readiness for deployment in real-world clinical settings.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Deep Learning, Citra Fundus, Multimodal, cWGAN-GP, SMOTE, Klasifikasi Penyakit Mata, Computer-Aided Diagnosis. |
| Subjects: | T Technology > TA Engineering (General). Civil engineering (General) T Technology > TA Engineering (General). Civil engineering (General) > TA1637 Image processing--Digital techniques. Image analysis--Data processing. |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 57201-(S1) Undergraduate Thesis |
| Depositing User: | Febrian Abdan Husnan |
| Date Deposited: | 27 Jul 2026 06:18 |
| Last Modified: | 27 Jul 2026 06:18 |
| URI: | http://repository.its.ac.id/id/eprint/137650 |
Actions (login required)
![]() |
View Item |
