Fadyah, Naila Syalwa (2026) Perbandingan K-means, K-modes, dan K-PbC untuk Klasterisasi Data Rumah Tangga di Jawa Timur. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5002221090-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (3MB) | Request a copy |
Abstract
Pengelompokan rumah tangga berdasarkan karakteristik sosial ekonomi merupakan salah satu pendekatan yang dapat digunakan untuk memahami heterogenitas kondisi masyarakat serta mendukung penyusunan kebijakan yang lebih tepat sasaran. Berbagai algoritma klasterisasi memiliki karakteristik yang berbeda, sehingga diperlukan penelitian untuk membandingkan kinerjanya. Penelitian ini bertujuan membandingkan kinerja algoritma K-means, K-modes, dan K-PbC dalam mengelompokkan rumah tangga berdasarkan karakteristik sosial ekonomi menggunakan dua skenario pemilihan fitur. Data yang digunakan merupakan data Survei Sosial Ekonomi Nasional (SUSENAS) tahun 2022 yang dibatasi pada kepala rumah tangga sebagai objek analisis. Tahapan penelitian meliputi prapemrosesan data, feature engineering, penyusunan dua skenario fitur, transformasi one-hot encoding untuk K-means, proses klasterisasi, evaluasi internal menggunakan Silhouette Score dan Davies-Bouldin Index (DBI), serta analisis stabilitas menggunakan Adjusted Rand Index (ARI) dan Normalized Mutual Information (NMI). Hasil penelitian menunjukkan bahwa hasil klasterisasi terbaik pada Skenario A diperoleh menggunakan algoritma K-means dengan k=3, menghasilkan Silhouette Score sebesar 0,3645 dan DBI sebesar 1,6726 serta memiliki kestabilan terbaik berdasarkan nilai ARI dan NMI. Sementara itu, hasil terbaik pada Skenario B diperoleh menggunakan algoritma K-PbC dengan k=6, menghasilkan Silhouette Score sebesar 0,3865 dan DBI sebesar 1,1161 dengan tingkat kestabilan yang baik. Di sisi lain, algoritma K-modes menghasilkan nilai Silhouette Score tertinggi sebesar 0,5461 dan DBI terendah sebesar 0,7284 pada Skenario A dengan k=10, tetapi memiliki tingkat kestabilan paling rendah dibandingkan algoritma lainnya. Selain menghasilkan karakteristik klaster yang berbeda pada setiap algoritma, hasil klasterisasi juga dapat dimanfaatkan sebagai informasi pendukung dalam penentuan sasaran program prioritas Jatim Sejahtera di Provinsi Jawa Timur, khususnya pada program perlindungan sosial, pemberdayaan ekonomi, dan transformasi digital rumah tangga.
==============================================================================================================================
Household clustering based on socioeconomic characteristics is one approach that can be used to understand the heterogeneity of community conditions and support the formulation of more targeted public policies. Various clustering algorithms have different characteristics, making it necessary to compare their performance. This study aims to compare the performance of the K-means, K-modes, and K-PbC algorithms in clustering households based on socioeconomic characteristics using two feature selection scenarios. The data used were obtained from the 2022 National Socioeconomic Survey (SUSENAS), with the analysis limited to household heads as the object of analysis. The research stages included data preprocessing, feature engineering, the development of two feature selection scenarios, one-hot encoding transformation for K-means, the clustering process, internal evaluation using the Silhouette Score and the Davies-Bouldin Index (DBI), and stability analysis using the Adjusted Rand Index (ARI) and the Normalized Mutual Information (NMI). The results showed that the best clustering performance for Scenario A was achieved by the K-means algorithm with k=3, producing a Silhouette Score of 0.3645 and a DBI of 1.6766, while also demonstrating the highest stability based on the ARI and NMI values. Meanwhile, the best performance for Scenario B was achieved by the K-PbC algorithm with k=6, producing a Silhouette Score of 0.3865 and a DBI of 1.1161 with good clustering stability. On the other hand, the K-modes algorithm achieved the highest Silhouette Score of 0.5461 and the lowest DBI of 0.7284 for Scenario A with k=10, but exhibited the lowest stability among the three algorithms. In addition to producing different cluster characteristics across algorithms, the clustering results can also serve as supporting information for determining the target beneficiaries of the Jatim Sejahtera priority programs in East Java Province, particularly in social protection, economic empowerment, and household digital transformation programs.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Clustering, Households, K-means, K-modes, K-PbC, Klasterisasi, Rumah Tangga, SUSENAS |
| Subjects: | Q Science > QA Mathematics > QA278.55 Cluster analysis Q Science > QA Mathematics > QA76.9.D343 Data mining. Querying (Computer science) Q Science > QA Mathematics > QA9.58 Algorithms |
| Divisions: | Faculty of Mathematics and Science > Mathematics > 44201-(S1) Undergraduate Thesis |
| Depositing User: | Naila Syalwa Fadyah |
| Date Deposited: | 29 Jul 2026 01:07 |
| Last Modified: | 29 Jul 2026 01:07 |
| URI: | http://repository.its.ac.id/id/eprint/139273 |
Actions (login required)
![]() |
View Item |
