Integrasi Data Multi-Omics Menggunakan Multi-Modal Neural Network untuk Klasifikasi Subtipe PAM50 dan Identifikasi Biomarker Kanker Payudara

Aurelia, Stephanie (2026) Integrasi Data Multi-Omics Menggunakan Multi-Modal Neural Network untuk Klasifikasi Subtipe PAM50 dan Identifikasi Biomarker Kanker Payudara. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5003221194-Undergraduate_Thesis.pdf] Text
5003221194-Undergraduate_Thesis.pdf - Accepted Version
Restricted to Repository staff only

Download (2MB) | Request a copy

Abstract

Tingginya insiden global dan kompleksitas subtipe molekuler kanker payudara menuntut identifikasi biomarker yang stabil untuk mendukung terapi presisi. Analisis terpisah pada satu modalitas omics tidak mampu menangkap interaksi multilevel biologis serta terkendala oleh dimensi data yang sangat tinggi (high-dimensional data). Penelitian ini bertujuan mengembangkan pipeline analisis integrasi komprehensif empat lapisan data multi-omics (Mutation, Methylation, RNA-Seq, dan RPPA) dari TCGA-BRCA dengan ukuran sampel teriris sebanyak 443 pasien untuk klasifikasi subtipe PAM50. Metodologi diawali dengan preprocessing penyelarasan ID pasien dan pembagian dataset train-test (80:20). Strategi reduksi dimensi membandingkan algoritma Boruta (supervised feature selection) dan Autoencoder dengan Pseudo-Huber Loss (unsupervised deep feature extraction), yang kinerjanya dievaluasi menggunakan algoritma Random Forest (RF) dan Multi-Modal Neural Network (MMNN) pada data pengujian (testing data). Hasil pengujian membuktikan pendekatan supervised selection jauh lebih unggul daripada unsupervised extraction. Skenario Autoencoder + Random Forest menghasilkan performa terendah dengan akurasi global 80,90% dan Macro-F1 Score 0,6773 akibat terbuangnya sinyal biologis penting selama kompresi ruang laten. Sebaliknya, skenario Boruta + Multi-Modal Neural Network menjadi model terbaik dengan capaian akurasi global tertinggi 87,64% dan Macro-F1 Score 0,8665, serta terbukti sangat robust mengenali subtipe minoritas ekstrem Her2-enriched. Analisis kontribusi fitur berbasis nilai SHAP (Shapley Additive exPlanations) berhasil mengidentifikasi molekuler spesifik penentu subtipe, di antaranya ekspresi gen CLGN yang rendah pada subtipe Basal-like namun tinggi pada Her2-enriched, penekanan ekspresi gen KAT2A pada Luminal A, serta ekspresi tinggi gen ANTXR2 pada Luminal B. Kesimpulannya, integrasi multi-omics berbasis seleksi fitur Boruta dan arsitektur deep learning MMNN mampu menghasilkan klasifikasi optimal dan interpretabilitas biomarker yang valid untuk mendukung onkologi presisi.
=========================================================================================================================================
The high global incidence and complexity of molecular breast cancer subtypes demand the identification of stable biomarkers to support precision therapy strategies. Conventional isolated analysis of a single omics modality fails to capture multilevel biological interactions and is heavily constrained by high-dimensional data. This study aims to develop a comprehensive analysis pipeline for integrating four multi-omics data layers (Mutation, Methylation, RNA-Seq, and RPPA) from the TCGA-BRCA dataset, with an intersected sample size of 443 patients, to classify PAM50 subtypes. The methodology begins with data preprocessing via patient ID alignment and a train-test dataset split (80:20). Dimensionality reduction strategies compare the Boruta algorithm (supervised feature selection) against an Autoencoder utilizing Pseudo-Huber Loss (unsupervised deep feature extraction), with their performances evaluated using Random Forest (RF) and Multi-Modal Neural Network (MMNN) algorithms on independent testing data. The evaluation results demonstrate that the supervised selection approach significantly outperforms unsupervised extraction. The Autoencoder + Random Forest scenario yields the lowest performance, with a global accuracy of 80.90% and a Macro-F1 Score of 0.6773, due to the loss of critical biological signals during latent space compression. Conversely, the Boruta + Multi-Modal Neural Network scenario achieves the best model performance, securing the highest global accuracy of 87.64% and a Macro-F1 Score of 0.8665, while proving highly robust in recognizing the extreme minority Her2-enriched subtype. Feature contribution analysis based on SHAP (Shapley Additive exPlanations) values successfully identifies subtype-specific molecular markers, including low expression of the CLGN gene in the Basal-like subtype but high in Her2-enriched, suppression of the KAT2A gene in Luminal A, and high expression of the ANTXR2 gene in Luminal B. In conclusion, multi-omics integration based on Boruta feature selection and the MMNN deep learning architecture provides optimal classification performance and clinically valid biomarker interpretability to support precision oncology.

Item Type: Thesis (Other)
Uncontrolled Keywords: Biomarker, Integrasi, Kanker Payudara, Klasifikasi, Multi-Omics, Biomarkers, Breast Cancer, Classification, Integration, Multi-Omics
Subjects: Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
Q Science > QA Mathematics > QA76.9D338 Data integration
R Medicine > R Medicine (General) > R858 Deep Learning
R Medicine > RC Internal medicine > RC0254 Neoplasms. Tumors. Oncology (including Cancer)
Divisions: Faculty of Science and Data Analytics (SCIENTICS) > Statistics > 49201-(S1) Undergraduate Thesis
Depositing User: Stephanie Aurelia
Date Deposited: 05 Aug 2026 03:54
Last Modified: 05 Aug 2026 03:54
URI: http://repository.its.ac.id/id/eprint/143938

Actions (login required)

View Item View Item