Putra, Davin Fisabilillah Reynard (2026) Analisis Efektivitas Teknik Resampling Dalam Penanganan Imbalanced Data Untuk Meningkatkan Sensitivitas Model Ensemble Learning Pada Dataset Penipuan Dalam Transaksi Kartu Kredit. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221137-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (11MB) | Request a copy |
Abstract
Deteksi fraud pada transaksi kartu kredit merupakan permasalahan penting dalam machine learning karena karakteristik datanya yang sangat tidak seimbang (imbalanced data), sehingga model tidak dapat mempelajari dengan baik pola transaksi fraud dikarenakan volume datanya yang sangat sedikit. Penelitian ini bertujuan menganalisis efektivitas berbagai teknik resampling dalam menangani imbalanced data untuk meningkatkan sensitivitas model ensemble learning pada kasus deteksi fraud transaksi kartu kredit. Penelitian menggunakan dataset Credit Card Fraud Detection yang terdiri dari 284.807 transaksi dengan proporsi fraud sebesar 0,172%. Tahapan penelitian meliputi data preprocessing, penerapan delapan teknik resampling yaitu Random Oversampling, Random Undersampling, SMOTE, ADASYN, Tomek Links, Edited Nearest Neighbor (ENN), SMOTE-ENN, dan SMOTE-Tomek, serta pelatihan model Random Forest, Gradient Boosting, XGBoost, Soft Voting, dan Stacking. Evaluasi dilakukan menggunakan metrik recall, precision, pr-auc, dan F1-Score dengan recall sebagai metrik utama. Selain itu, dilakukan hyperparameter tuning menggunakan RandomizedSearchCV pada lima kombinasi terbaik berdasarkan F1-Score. Hasil penelitian menunjukkan bahwa efektivitas teknik resampling berbeda pada setiap model. Teknik ENN menghasilkan rata-rata F1-Score tertinggi dengan nilai 0,8314. Untuk tujuan fraud detection yang memprioritaskan recall atau performa model dalam mendeteksi transaksi fraud, kombinasi XGBoost dan ENN menjadi yang paling efektif dengan recall sebesar 0,8367, precision senilai 0,9111, pr-auc sebesar 0,8517, dan F1-Score dengan nilai 0,8723. Penelitian ini menyimpulkan bahwa teknik pembersihan data seperti ENN dan Tomek Links lebih efektif dibandingkan teknik yang menyeimbangkan distribusi kelas yang agresif dalam meningkatkan kinerja sistem deteksi fraud.
====================================================================================================================================
Credit card fraud detection is a major challenge in machine learning due to the highly imbalanced nature of the transaction data, where fraudulent transactions represent only a small portion of all observations. The imbalance of the data volume between classes causes the models to not properly learn the patterns and characteristics of fraudulent transactions due to the limited amount of data. This study aims to analyze the effectiveness of various resampling techniques in handling imbalanced data to improve the sensitivity of ensemble learning models for credit card fraud detection. The research used the Credit Card Fraud Detection dataset consisting of 284,807 transactions, with fraudulent transactions accounting for only 0.172% of the data. The methodology included data preprocessing, the application of eight resampling techniques, namely Random Oversampling, Random Undersampling, SMOTE, ADASYN, Tomek Links, Edited Nearest Neighbor (ENN), SMOTE-ENN, and SMOTE-Tomek, followed by the training of Random Forest, Gradient Boosting, XGBoost, and Stacking models. Model performance was evaluated using Recall, Precision, pr-auc, and F1-Score, with Recall serving as the primary metric. Hyperparameter tuning using RandomizedSearchCV was also conducted on the top five model-resampling combinations. The results show that the effectiveness of resampling techniques varies across models. ENN achieved the highest average F1-Score with the score of 0.8314. From a fraud detection perspective where recall is prioritized because it measures the model’s performance based on the ability to detect fraudulent transactions, the XGBoost and ENN combination proved to be the most effective, achieving a Recall of 0.8367, Precision of 0.9111, PR-AUC of 0,8517, and F1-Score of 0.8723. The study concludes that data cleaning techniques such as ENN and Tomek Links are more effective than aggressive class-balancing approaches in improving fraud detection performance.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Fraud Detection, Imbalanced Data, Resampling, Ensemble Learning, XGBoost, Edited Nearest Neighbor, Recall. |
| Subjects: | Q Science > Q Science (General) > Q180.55.M38 Mathematical models Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. Q Science > QA Mathematics > QA336 Artificial Intelligence Q Science > QA Mathematics > QA401 Mathematical models. T Technology > T Technology (General) > T57.5 Data Processing T Technology > T Technology (General) > T57.8 Nonlinear programming. Support vector machine. Wavelets. Hidden Markov models. T Technology > T Technology (General) > T57.84 Heuristic algorithms. |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Davin Fisabilillah Reynard Putra |
| Date Deposited: | 25 Jul 2026 08:41 |
| Last Modified: | 25 Jul 2026 08:41 |
| URI: | http://repository.its.ac.id/id/eprint/137289 |
Actions (login required)
![]() |
View Item |
