Analisis Efektivitas Teknik Resampling Dalam Penanganan Imbalanced Data Untuk Meningkatkan Sensitivitas Model Ensemble Learning Pada Dataset Penipuan Dalam Transaksi Kartu Kredit

Putra, Davin Fisabilillah Reynard (2026) Analisis Efektivitas Teknik Resampling Dalam Penanganan Imbalanced Data Untuk Meningkatkan Sensitivitas Model Ensemble Learning Pada Dataset Penipuan Dalam Transaksi Kartu Kredit. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221137-Undergraduate_Thesis.pdf] Text
5025221137-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (11MB) | Request a copy

Abstract

Deteksi fraud pada transaksi kartu kredit merupakan permasalahan penting dalam machine learning karena karakteristik datanya yang sangat tidak seimbang (imbalanced data), sehingga model tidak dapat mempelajari dengan baik pola transaksi fraud dikarenakan volume datanya yang sangat sedikit. Penelitian ini bertujuan menganalisis efektivitas berbagai teknik resampling dalam menangani imbalanced data untuk meningkatkan sensitivitas model ensemble learning pada kasus deteksi fraud transaksi kartu kredit. Penelitian menggunakan dataset Credit Card Fraud Detection yang terdiri dari 284.807 transaksi dengan proporsi fraud sebesar 0,172%. Tahapan penelitian meliputi data preprocessing, penerapan delapan teknik resampling yaitu Random Oversampling, Random Undersampling, SMOTE, ADASYN, Tomek Links, Edited Nearest Neighbor (ENN), SMOTE-ENN, dan SMOTE-Tomek, serta pelatihan model Random Forest, Gradient Boosting, XGBoost, Soft Voting, dan Stacking. Evaluasi dilakukan menggunakan metrik recall, precision, pr-auc, dan F1-Score dengan recall sebagai metrik utama. Selain itu, dilakukan hyperparameter tuning menggunakan RandomizedSearchCV pada lima kombinasi terbaik berdasarkan F1-Score. Hasil penelitian menunjukkan bahwa efektivitas teknik resampling berbeda pada setiap model. Teknik ENN menghasilkan rata-rata F1-Score tertinggi dengan nilai 0,8314. Untuk tujuan fraud detection yang memprioritaskan recall atau performa model dalam mendeteksi transaksi fraud, kombinasi XGBoost dan ENN menjadi yang paling efektif dengan recall sebesar 0,8367, precision senilai 0,9111, pr-auc sebesar 0,8517, dan F1-Score dengan nilai 0,8723. Penelitian ini menyimpulkan bahwa teknik pembersihan data seperti ENN dan Tomek Links lebih efektif dibandingkan teknik yang menyeimbangkan distribusi kelas yang agresif dalam meningkatkan kinerja sistem deteksi fraud.
====================================================================================================================================
Credit card fraud detection is a major challenge in machine learning due to the highly imbalanced nature of the transaction data, where fraudulent transactions represent only a small portion of all observations. The imbalance of the data volume between classes causes the models to not properly learn the patterns and characteristics of fraudulent transactions due to the limited amount of data. This study aims to analyze the effectiveness of various resampling techniques in handling imbalanced data to improve the sensitivity of ensemble learning models for credit card fraud detection. The research used the Credit Card Fraud Detection dataset consisting of 284,807 transactions, with fraudulent transactions accounting for only 0.172% of the data. The methodology included data preprocessing, the application of eight resampling techniques, namely Random Oversampling, Random Undersampling, SMOTE, ADASYN, Tomek Links, Edited Nearest Neighbor (ENN), SMOTE-ENN, and SMOTE-Tomek, followed by the training of Random Forest, Gradient Boosting, XGBoost, and Stacking models. Model performance was evaluated using Recall, Precision, pr-auc, and F1-Score, with Recall serving as the primary metric. Hyperparameter tuning using RandomizedSearchCV was also conducted on the top five model-resampling combinations. The results show that the effectiveness of resampling techniques varies across models. ENN achieved the highest average F1-Score with the score of 0.8314. From a fraud detection perspective where recall is prioritized because it measures the model’s performance based on the ability to detect fraudulent transactions, the XGBoost and ENN combination proved to be the most effective, achieving a Recall of 0.8367, Precision of 0.9111, PR-AUC of 0,8517, and F1-Score of 0.8723. The study concludes that data cleaning techniques such as ENN and Tomek Links are more effective than aggressive class-balancing approaches in improving fraud detection performance.

Item Type: Thesis (Other)
Uncontrolled Keywords: Fraud Detection, Imbalanced Data, Resampling, Ensemble Learning, XGBoost, Edited Nearest Neighbor, Recall.
Subjects: Q Science > Q Science (General) > Q180.55.M38 Mathematical models
Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines.
Q Science > QA Mathematics > QA336 Artificial Intelligence
Q Science > QA Mathematics > QA401 Mathematical models.
T Technology > T Technology (General) > T57.5 Data Processing
T Technology > T Technology (General) > T57.8 Nonlinear programming. Support vector machine. Wavelets. Hidden Markov models.
T Technology > T Technology (General) > T57.84 Heuristic algorithms.
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Davin Fisabilillah Reynard Putra
Date Deposited: 25 Jul 2026 08:41
Last Modified: 25 Jul 2026 08:41
URI: http://repository.its.ac.id/id/eprint/137289

Actions (login required)

View Item View Item