Asy'ari, Zulchair (2026) Klasifikasi Durasi Rawat Inap Pasien Menggunakan Stacking Ensemble Berbasis Tabular Deep Learning Dengan Pendekatan Explainable AI. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6025222005_Buku Thesis_Final_stempel.pdf - Accepted Version Restricted to Repository staff only Download (1MB) | Request a copy |
Abstract
Durasi rawat inap pasien (Length of Stay/LOS) merupakan indikator penting dalam manajemen rumah sakit karena memengaruhi efisiensi pelayanan, pemanfaatan sumber daya, dan kualitas layanan. Prediksi durasi rawat inap yang akurat dapat membantu rumah sakit mengantisipasi kebutuhan kapasitas serta mendukung pengambilan keputusan klinis dan manajerial. Penelitian ini mengusulkan model klasifikasi durasi rawat inap dengan pendekatan ensemble learning berbasis stacking yang mengombinasikan TabNet, CatBoost, dan FT-Transformer. Ketiga model dipilih karena memiliki karakteristik yang saling melengkapi, yaitu kemampuan CatBoost menangani fitur kategorikal, TabNet dengan seleksi fitur adaptif, serta FT-Transformer dalam memodelkan interaksi antar fitur melalui mekanisme self-attention. Dataset yang digunakan berasal dari data administrasi pasien rumah sakit periode 2019–2025 dengan total 671.915 data dan 20 fitur. Variabel LOS diklasifikasikan ke dalam tiga kategori, yaitu durasi pendek (≤3 hari), sedang (4–7 hari), dan panjang (≥8 hari). Proses pengembangan model meliputi tahap praproses, optimasi hyperparameter dengan Optuna, serta pembangunan stacking ensemble dengan memanfaatkan prediksi probabilitas out-of-fold sebagai input bagi Logistic Regression sebagai meta-learner. Hasil penelitian menunjukkan bahwa pendekatan ensemble stacking mencapai akurasi sebesar 81,15% dan weighted F1-score sebesar 81,22% serta meningkatkan kemampuan identifikasi pasien durasi panjang dibandingkan seluruh model tunggal, termasuk CatBoost sebagai model tunggal dengan performa terbaik. Analisis SHAP dan LIME menunjukkan bahwa NamaRuangan, CaraMasuk, dan NamaDiagnosa memberikan kontribusi penting terhadap prediksi LOS serta meningkatkan interpretabilitas dan transparansi model. Model yang diusulkan menunjukkan bahwa penggabungan model tabular deep learning melalui stacking mampu meningkatkan performa klasifikasi sekaligus menyediakan interpretasi yang transparan terhadap hasil prediksi, sehingga berpotensi mendukung manajemen sumber daya dan pengambilan keputusan berbasis data di rumah sakit.
===================================================================================================================================
Hospital Length of Stay (LOS) is a key indicator in hospital management because it affects service efficiency, resource utilization, and the quality of healthcare services. Accurate LOS prediction can help hospitals anticipate capacity requirements and support both clinical and managerial decision-making. This study proposes an inpatient LOS classification model using a stacking-based ensemble learning approach that integrates TabNet, CatBoost, and FT-Transformer. These three models were selected because of their complementary learning characteristics: CatBoost effectively handles categorical features, TabNet performs adaptive feature selection, and FT-Transformer models complex feature interactions through a self-attention mechanism. The dataset consists of hospital administrative records collected between 2019 and 2025, comprising 671,915 samples with 20 input features. The LOS variable is categorized into three classes: Short Stay (≤3 days), Mid Stay (4–7 days), and Long Stay (≥8 days). The proposed framework includes data preprocessing, hyperparameter optimization using Optuna, and the construction of a stacking ensemble that utilizes out-of-fold probability predictions as input to a Logistic Regression meta-learner. The experimental results demonstrate that the proposed stacking ensemble achieved an accuracy of 81.15% and a weighted F1-score of 81.22%, while improving the identification of Long Stay patients compared with all individual models, including CatBoost, which achieved the best performance among the standalone models. SHapley Additive exPlanations (SHAP) and Local Interpretable Model-Agnostic Explanations (LIME) analyses reveal that NamaRuangan, CaraMasuk, and NamaDiagnosa are among the most influential features contributing to LOS prediction, thereby enhancing the interpretability and transparency of the proposed model. The findings indicate that integrating tabular deep learning models through stacking can improve classification performance while providing transparent explanations of the prediction process, making the proposed framework a promising tool for supporting hospital resource management and data-driven decision-making.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | CatBoost, Deep Learning, durasi rawat inap, Ensemble Learning, FT Transformer, Gradient Boosting, Klasifikasi, LIME, SHAP, TabNet, XAI, Classification, Length of Stay |
| Subjects: | Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. Q Science > Q Science (General) > Q337.5 Pattern recognition systems |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis |
| Depositing User: | Zulchair Asy'ari |
| Date Deposited: | 04 Aug 2026 03:45 |
| Last Modified: | 04 Aug 2026 03:45 |
| URI: | http://repository.its.ac.id/id/eprint/142882 |
Actions (login required)
![]() |
View Item |
