Perbandingan Algoritma Random Forest dan XGBoost dalam Klasifikasi Pemanfaatan Fasilitas Kesehatan Rujukan Tingkat Lanjut (FKRTL) Peserta BPJS Kesehatan Jawa Timur

Naja, Fadila Shahnun (2026) Perbandingan Algoritma Random Forest dan XGBoost dalam Klasifikasi Pemanfaatan Fasilitas Kesehatan Rujukan Tingkat Lanjut (FKRTL) Peserta BPJS Kesehatan Jawa Timur. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5006221052-Undergraduate_Thesis.pdf] Text
5006221052-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (4MB) | Request a copy

Abstract

Program Jaminan Kesehatan Nasional (JKN) yang dikelola BPJS Kesehatan menghadapi tantangan keberlanjutan pembiayaan akibat tingginya rasio klaim yang mencapai 107,93% pada kuartal II tahun 2024, di mana tingginya angka rujukan peserta ke Fasilitas Kesehatan Rujukan Tingkat Lanjut (FKRTL) menjadi salah satu penyebab utama defisit anggaran. Penelitian ini bertujuan membandingkan kinerja algoritma Random Forest dan XGBoost dalam mengklasifikasikan pemanfaatan FKRTL oleh peserta BPJS Kesehatan di Provinsi Jawa Timur, serta menganalisis kontribusi masing-masing variabel melalui feature importance. Random Forest dan XGBoost dipilih karena keunggulannya dalam menangani data kompleks dengan tingkat akurasi tinggi. Random Forest menerapkan pendekatan bagging yang mampu mengurangi overfitting dengan membangun banyak pohon keputusan secara independen, sedangkan XGBoost menggunakan pendekatan boosting yang memperbaiki kesalahan secara berurutan sehingga menghasilkan model yang lebih akurat dan efisien. Penelitian difokuskan pada Provinsi Jawa Timur mengingat cakupan kepesertaan JKN mencapai 95,98% atau 40.038.331 jiwa pada Desember 2024. Data yang digunakan merupakan data sampel BPJS Kesehatan tahun 2024 sebanyak 526.003 kunjungan peserta dengan variabel usia, jenis kelamin, status pernikahan, segmentasi peserta, frekuensi kedatangan, dan diagnosis berdasarkan klasifikasi ICD-10. Ketidakseimbangan kelas dengan rasio 11:1 diatasi menggunakan SMOTE-NC, dan optimasi hyperparameter dilakukan menggunakan Optuna dengan 10-fold cross-validation sebanyak 50 trial. Hasil evaluasi menunjukkan kedua algoritma memiliki performa yang hampir sama dengan Random Forest menghasilkan akurasi 85,61%, F1-Score 44,19%, dan AUC 0,88, sedangkan XGBoost menghasilkan akurasi 85,62%, F1-Score 44,71%, dan AUC 0,87. XGBoost dipilih sebagai algoritma lebih optimal karena unggul pada metrik akurasi, presisi, recall, F1-Score, serta efisiensi komputasi. Analisis feature importance XGBoost menunjukkan diagnosis sebagai variabel paling berkontribusi, dengan diagnosis congenital malformations and deformations memiliki persentase rujukan tertinggi. Hasil penelitian ini diharapkan dapat membantu BPJS Kesehatan dalam mengidentifikasi pola pemanfaatan FKRTL dan mendukung pengambilan keputusan berbasis data untuk pengelolaan anggaran layanan rujukan.
=================================================================================================================================================
The National Health Insurance (JKN) program managed by BPJS Kesehatan faces sustainability challenges due to a high claims ratio reaching 107.93% in the second quarter of 2024, where the high referral rate of participants to Advanced Referral Health Facilities (FKRTL) is one of the main contributors to the budget deficit. This study aims to compare the performance of Random Forest and XGBoost algorithms in classifying FKRTL utilization by BPJS Kesehatan participants in East Java Province, as well as to analyze the contribution of each variable through feature importance. Random Forest and XGBoost were selected due to their advantages in handling complex data with high accuracy. Random Forest applies a bagging approach that reduces overfitting by building many decision trees independently, while XGBoost uses a boosting approach that corrects errors sequentially, resulting in a more accurate and efficient model. The study focuses on East Java Province given that JKN membership coverage in the region reached 95.98%, or 40,038,331 participants as of December 2024. The data used is a BPJS Kesehatan sample dataset from 2024 covering 526,003 participant visits with variables including age, gender, marital status, participant segmentation, visit frequency, and diagnosis based on ICD-10 classification. Class imbalance with a ratio of 11:1 was addressed using SMOTE-NC, and hyperparameter optimization was performed using Optuna with 10-fold cross-validation over 50 trials. The evaluation results show that both algorithms achieved comparable performance, with Random Forest yielding an accuracy of 85.61%, F1 Score of 44.19%, and AUC of 0.88, while XGBoost yielded an accuracy of 85.62%, F1-Score of 44.71%, and AUC of 0.87. XGBoost was selected as the more optimal algorithm as it outperformed in accuracy, precision, recall, F1-Score, and computational efficiency. Feature importance analysis from XGBoost revealed diagnosis as the most contributing variable, with the congenital malformations and deformations category having the highest referral percentage. The findings of this study are expected to assist BPJS Kesehatan in identifying FKRTL utilization patterns and support data-driven decision-making for referral service budget management.

Item Type: Thesis (Other)
Uncontrolled Keywords: BPJS Kesehatan, Machine Learning, Optuna, Random Forest, XGBoost.
Subjects: Q Science > QA Mathematics > QA76.6 Computer programming.
Q Science > QA Mathematics > QA76.9D338 Data integration
Q Science > QA Mathematics > QA9.58 Algorithms
Divisions: Faculty of Science and Data Analytics (SCIENTICS) > Actuaria > 94203-(S1) Undergraduate Thesis
Depositing User: Fadila Shahnun Naja
Date Deposited: 17 Jul 2026 07:18
Last Modified: 17 Jul 2026 07:18
URI: http://repository.its.ac.id/id/eprint/135348

Actions (login required)

View Item View Item