Deteksi Outlier Berbasis Klaster pada Set Data Dengan Atribut Campuran Numerik dan Kategorikal

Maryono, Dwi (2010) Deteksi Outlier Berbasis Klaster pada Set Data Dengan Atribut Campuran Numerik dan Kategorikal. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5107201006-Master_Thesis.pdf] Text
5107201006-Master_Thesis.pdf
Restricted to Repository staff only

Download (18MB) | Request a copy

Abstract

Deteksi outlier merupakan salah satu bidang penelitian yang penting dalam topik data mining. Penelitian ini sangat bermanfaat untuk mendeteksi perilaku yang tidak normal, seperti deteksi penipuan menggunakan kartu kredit, deteksi intrusi jaringan, diagnosis medis, dan lain-lain. Ada banyak metode yang dikembangkan, baik dari segi teknik maupun objek yang diteliti. Namun, di antara sekian banyak metode tersebut terjadi dikotomi, bahwa metode-metode tersebut hanya fokus pada data dengan atribut yang seragam, yaitu data numerik saja atau data kategorikal saja. Di sisi lain, data di lapangan sering kali merupakan gabungan dari dua jenis atribut tersebut. Penelitian data mining mengenai set data dengan atribut campuran seperti ini dapat dikatakan masih sangat jarang. Dalam penelitian ini diajukan sebuah metode, yaitu MixCBLOF, untuk mendeteksi adanya outlier pada set data campuran seperti ini. Pendekatan yang digunakan untuk menyelesaikan masalah ini adalah gabungan dari beberapa teknik, seperti klasterisasi subdata, deteksi outlier berbasis klaster pada subdata numerik dan kategorikal, serta penggunaan Multi-Attribute Decision Making (MADM) untuk penggabungan derajat akhir dari setiap objek data. Evaluasi dilakukan pada beberapa set data nyata yang diperoleh dari UCI Machine Learning Repository. Evaluasi dilakukan dengan membandingkan rata-rata coverage untuk top ratio antara jumlah outlier eksak dengan keseluruhan data yang dievaluasi. Dari uji coba yang dilakukan, diperoleh hasil bahwa algoritma MixCBLOF cukup efektif untuk mendeteksi outlier pada set data campuran dengan rata-rata pencapaian coverage sebesar 73,54%. Hasil ini lebih baik jika dibandingkan dengan algoritma CBLOF yang diterapkan pada set data yang sama yang telah didiskritisasi dengan metode equal width, yang hanya menghasilkan rata-rata coverage sebesar 59,48%.
==================================================================================================================================
Outlier detection is one of the important research areas in data mining. This research is very beneficial for detecting abnormal behaviors, such as credit card fraud, network intrusion, medical diagnosis, and so on. Many methods have been developed, either in terms of techniques or the objects being studied. However, among the methods developed, there is a dichotomy in which the methods only focus on data containing uniform attributes, i.e., numerical or categorical data. On the other hand, data often contain a combination of both numerical and categorical attributes. Research on outlier detection involving datasets with mixed attributes is relatively rare. Therefore, in this research, a modified Cluster-Based Local Outlier Factor (CBLOF) algorithm was proposed. The algorithm, called MixCBLOF, was specifically designed to detect outliers in mixed datasets. The MixCBLOF algorithm involves several techniques, such as subdata clustering, cluster-based outlier detection, and the use of Multi-Attribute Decision Making (MADM) to integrate the final outlier factors of data objects. The proposed method was evaluated using several datasets obtained from the UCI Machine Learning Repository. The evaluation was performed by comparing the average coverage for the top ratio between the number of exact outliers and the total number of data being evaluated. Experimental results show that the MixCBLOF algorithm is sufficiently effective in detecting outliers in mixed datasets, with an average coverage of 73.54%. This result was better than that produced by the CBLOF algorithm applied to the same datasets discretized using the equal-width approach, which achieved an average coverage of 59.48%.

Item Type: Thesis (Masters)
Additional Information: RTIf 006.312 Mar d
Uncontrolled Keywords: data campuran, deteksi outlier, outlier berbasis klaster, CBLOF, mixed dataset, outlier detection, cluster-based outlier, CBLOF
Subjects: Q Science > QA Mathematics > QA76.9.D343 Data mining. Querying (Computer science)
Divisions: Faculty of Information Technology > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: magang .
Date Deposited: 21 Sep 2026 03:08
Last Modified: 21 Sep 2026 03:08
URI: http://repository.its.ac.id/id/eprint/144719

Actions (login required)

View Item View Item