Pengembangan Model Unsupervised Feature Engineering Menggunakan Deep Learning Dengan Contrastive Learning Untuk Data Tabular

Hermafidhanti, Fairna Mustika (2026) Pengembangan Model Unsupervised Feature Engineering Menggunakan Deep Learning Dengan Contrastive Learning Untuk Data Tabular. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5025221160-Undergraduate_Thesis.pdf] Text
5025221160-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (8MB) | Request a copy

Abstract

Data tabular merupakan bentuk data yang paling umum digunakan di berbagai sektor krusial seperti kesehatan, keuangan, dan industri. Pola informasi pada data tabular bersifat kompleks dan nonlinier sehingga membutuhkan representasi fitur yang tepat untuk dapat dianalisis secara efektif. Pada praktiknya, data tabular didominasi oleh data tidak berlabel akibat mahalnya proses pelabelan, sehingga pendekatan deep learning konvensional rentan terhadap overfitting dan ketidakstabilan pembelajaran. Kondisi ini mendorong perlunya strategi pembentukan representasi fitur yang informatif tanpa ketergantungan pada label.
Penelitian ini mengembangkan metode unsupervised contrastive learning dua fase sebagai feature engineering untuk data tabular. Training Fase I melatih encoder dengan margin contrastive loss berbasis augmentasi marginal corruption. Training Fase II memperkuat encoder dengan margin contrastive loss berbasis pseudo-label K-Means dari embedding Fase I. Evaluasi dilakukan pada lima dataset menggunakan enam algoritma clustering dan empat metrik yaitu Silhouette Score, Davies-Bouldin Index, Dunn Index, dan Adjusted Rand Index. Hasil eksperimen menunjukkan metode usulan menjadi metode tertinggi pada Silhouette Score, Davies-Bouldin Index, dan Adjusted Rand Index. Keunggulan ini paling konsisten pada tiga dari lima dataset yang diuji. Capaian metode usulan melebihi seluruh metode pembanding termasuk fitur asli, AE, DAE, VAE, SCARF, dan Polynomial. Metode usulan juga menunjukkan performa terbaik pada dataset dengan C-Score rendah. Hasil penelitian menunjukkan bahwa metode yang diusulkan mampu menghasilkan representasi fitur yang lebih informatif sehingga meningkatkan kualitas clustering pada data tabular tanpa label.
========================================================================================================================
==========
Tabular data is the most commonly used data format in various critical sectors such as healthcare, finance, and industry. Information patterns in tabular data are complex and nonlinear, requiring appropriate feature representation to enable effective analysis. In practice, tabular data is dominated by unlabeled data due to the costly labeling process, causing conventional deep learning approaches to be prone to overfitting and learning instability. This condition drives the need for strategies to form informative feature representations without dependency on labels.
This research develops a two-phase unsupervised contrastive learning method as feature engineering for tabular data. Training Phase I trains the encoder with margin contrastive loss based on marginal corruption augmentation. Training Phase II refines the encoder with margin contrastive loss based on K-Means pseudo-labels derived from Phase I embeddings. Evaluation was conducted on five datasets using six clustering algorithms and four metrics, namely Silhouette Score, Davies-Bouldin Index, Dunn Index, and Adjusted Rand Index. Experimental results show that the proposed method achieves the highest performance on Silhouette Score, Davies-Bouldin Index, and Adjusted Rand Index. This superiority is most consistent on three out of five tested datasets. The proposed method outperforms all comparison methods including original features, AE, DAE, VAE, SCARF, and Polynomial. The proposed method also achieved the best performance on datasets with low C-Scores, demonstrating its ability to learn more informative feature representations and effectively improve clustering quality for unlabeled tabular data.

Item Type: Thesis (Other)
Uncontrolled Keywords: Contrastive Learning, Data Tabular, Deep Learning, Feature Engineering, Unsupervised Learning, Contrastive Learning, Deep Learning, Feature Engineering, Tabular Data, Unsupervised Learning
Subjects: Q Science > QA Mathematics > QA278.55 Cluster analysis
Q Science > QA Mathematics > QA336 Artificial Intelligence
T Technology > T Technology (General) > T57.5 Data Processing
Divisions: Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis
Depositing User: Fairna Mustika Hermafidhanti
Date Deposited: 24 Jul 2026 05:59
Last Modified: 24 Jul 2026 05:59
URI: http://repository.its.ac.id/id/eprint/137131

Actions (login required)

View Item View Item