Nafis, Raka Salman (2026) Perbandingan Metode Data Balancing pada Algoritma Machine Learning untuk Pemodelan Risiko Penjualan Kredit Perusahaan dengan Interpretabilitas Berbasis SHAP. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5006221029-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (4MB) | Request a copy |
Abstract
Penjualan kredit pada perusahaan Fast Moving Consumer Goods (FMCG) merupakan strategi penting untuk meningkatkan volume penjualan, namun memiliki risiko tinggi terhadap terjadinya gagal bayar yang umumnya diawali oleh keterlambatan pembayaran (overdue) secara berulang pada sebagian pelanggan, dengan distribusi data kejadian yang cenderung tidak seimbang (imbalanced data). Kondisi ini menyebabkan model klasifikasi machine learning tanpa penanganan khusus menjadi bias terhadap kelas mayoritas, yang tercermin dari nilai accuracy global yang tinggi tetapi diikuti oleh nilai recall kelas overdue yang rendah. Penelitian ini bertujuan untuk membandingkan kinerja model klasifikasi serta mengidentifikasi kombinasi algoritma klasifikasi machine learning, metode data balancing dan subset fitur paling optimal dalam memprediksi kejadian overdue pada pelanggan FMCG. Data penjualan kredit diproses menggunakan metode Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic Sampling (ADASYN), dan Synthetic Minority Oversampling Technique dengan kombinasi Edited Nearest Neighbors (SMOTE-ENN), kemudian dimodelkan menggunakan algoritma Extreme Gradient Boosting (XGBoost) dan Light Gradient Boosting Machine (LightGBM). Hasil penelitian menunjukkan bahwa penerapan metode data balancing secara signifikan meningkatkan kemampuan model dalam mendeteksi kelas minoritas melalui peningkatan nilai recall, meskipun disertai trade-off berupa penurunan nilai accuracy dan precision. Kombinasi terbaik diperoleh pada model XGBoost dengan metode ADASYN yang dioptimalkan menggunakan hyperparameter tuning berbasis RandomSearchCV, dengan nilai recall sebesar 0,865, accuracy 0,618, precision 0,347, F1-score 0,495, dan AUC 0,806. Analisis feature importance berbasis SHapley Additive exPlanations (SHAP) menunjukkan bahwa variabel Jumlah Bounced, Credit Limit, dan Jumlah Invoice merupakan fitur paling berpengaruh. Selain itu, penggunaan subset enam fitur terpilih berdasarkan nilai kontribusi tertinggi SHAP terbukti mampu memberikan kinerja prediktif yang relatif lebih baik dibandingkan model dengan seluruh fitur data. Model yang dikembangkan diharapkan dapat mendukung pengambilan kebijakan kredit yang lebih tepat sasaran serta berperan sebagai sistem peringatan dini terhadap potensi keterlambatan pembayaran pelanggan.
=====================================================================================================================================
Credit sales in Fast Moving Consumer Goods (FMCG) companies are widely applied to increase sales volume. However, they also expose companies to a substantial risk of default, which is often preceded by repeated payment delays (overdue) among a small proportion of customers, resulting in highly imbalanced data distributions. Under such conditions, conventional machine learning classification models tend to be biased toward the majority class, producing high overall accuracy while exhibiting poor capability in detecting overdue cases. This study aims to compare the performance of classification models and identify the optimal combination of machine learning algorithms, data balancing techniques, and feature subsets for predicting overdue events among FMCG customers. Credit sales data were processed using the Synthetic Minority Oversampling Technique (SMOTE), Adaptive Synthetic Sampling (ADASYN), and SMOTE combined with Edited Nearest Neighbors (SMOTE-ENN), and subsequently modeled using Extreme Gradient Boosting (XGBoost) and Light Gradient Boosting Machine (LightGBM). The results indicate that data balancing techniques substantially improve the detection of the minority class by increasing recall, although this improvement comes at the expense of lower accuracy and precision. The best performing model was achieved using XGBoost with ADASYN and RandomSearchCV based hyperparameter tuning, attaining a recall of 0.865, an accuracy of 0.618, a precision of 0.347, an F1-score of 0.495, and an AUC of 0.806. Feature importance analysis based on SHapley Additive exPlanations (SHAP) identified Number of Bounced Payments, Credit Limit, and Number of Invoices as the most influential predictors. Moreover, a six-feature subset selected according to SHAP importance yielded slightly better predictive performance than the full-feature model. The proposed model provides a data-driven approach to supporting credit policy decisions and has the potential to serve as an effective early warning system for identifying customers at risk of payment delinquency.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Data Balancing, FMCG, LightGBM, Penjualan Kredit, SHAP, XGBoost, Credit Sales, Data Balancing, FMCG, LightGBM, SHAP, XGBoost. |
| Subjects: | H Social Sciences > HG Finance > HG3751 Credit--Management. Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. T Technology > T Technology (General) > T174.5 Technology--Risk assessment. |
| Divisions: | Faculty of Science and Data Analytics (SCIENTICS) > Actuaria > 94203-(S1) Undergraduate Thesis |
| Depositing User: | Raka Salman Nafis |
| Date Deposited: | 17 Jul 2026 07:33 |
| Last Modified: | 17 Jul 2026 07:33 |
| URI: | http://repository.its.ac.id/id/eprint/135334 |
Actions (login required)
![]() |
View Item |
