Pramana, Revy (2026) Analisis Komparatif Model Deteksi Anomali pada Log Web dengan Preprocessing Berbasis Urutan Pengaksesan Url di Aws Clean Rooms. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221252-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (11MB) | Request a copy |
Abstract
Log akses web merupakan rekam jejak digital untuk mengidentifikasi ancaman siber melalui pola permintaan URL yang tidak wajar. Namun, volume data log yang sangat masif dan tidak terstruktur menyulitkan proses analisis secara manual. Sebagai solusi, penerapan deteksi anomali berbasis unsupervised learning sering digunakan untuk mengenali pola serangan tanpa membutuhkan data berlabel. Meskipun demikian, upaya peningkatan kualitas deteksi melalui skema kolaborasi antarorganisasi sering kali terbentur oleh regulasi privasi yang melarang pertukaran data mentah secara bebas. Untuk mengatasi masalah tersebut, penelitian ini mengusulkan implementasi lingkungan kolaborasi data yang aman menggunakan arsitektur AWS Clean Rooms. Sistem ini dipadukan dengan preprocessing berbasis urutan pengaksesan URL menggunakan metode N-Gram untuk mengekstraksi informasi struktural secara efisien. Penelitian ini melakukan analisis komparatif terhadap empat model yang mencakup Isolation Forest, Local Outlier Factor, One-Class Support Vector Machine, dan Autoencoder. Berdasarkan pengujian yang dilakukan, skema kolaborasi AWS Clean Rooms sukses mengekspos model pada ragam ancaman tingkat lanjut yang jauh lebih kompleks dan realistis. Model Autoencoder berhasil menunjukkan tingkat adaptasi paling tangguh dalam menghadapi dinamika data gabungan dengan mencetak nilai ketepatan prediksi sebesar 0,8811 pada ekstraksi sekuens 4-Gram. Sementara itu, algoritma konvensional mengalami penurunan kinerja yang signifikan pada lingkungan kolaborasi akibat tingginya angka peringatan palsu. Penurunan performa tersebut dipicu oleh ketidakmampuan model konvensional dalam membedakan teks rute URL aset web statis yang panjang dengan pola serangan siber yang sesungguhnya. Di sisi lain, algoritma Isolation Forest mengalami kegagalan deteksi dengan mencetak nilai akurasi terendah di hampir seluruh skenario pengujian. Oleh karena itu, penggunaan arsitektur Autoencoder dengan ekstraksi fitur 4-Gram di dalam skema AWS Clean Rooms disimpulkan sebagai konfigurasi yang paling direkomendasikan.
===================================================================================================================================
Web access logs are digital footprints for identifying cyber threats through irregular URL request patterns. However, the massive volume and unstructured nature of log data complicate the manual analysis process. As a solution, unsupervised learning-based anomaly detection is frequently applied to recognize attack patterns without requiring labeled data. Nevertheless, efforts to enhance detection quality through inter-organizational collaboration schemes are often hindered by privacy regulations that prohibit the free exchange of raw data. To overcome this issue, this research proposes the implementation of a secure data collaboration environment using the AWS Clean Rooms architecture. This system is integrated with URL access sequence-based preprocessing using the N-Gram technique to efficiently extract structural information. This study conducts a comparative analysis of four models encompassing Isolation Forest, Local Outlier Factor, One-Class Support Vector Machine, and Autoencoder. Results from the conducted tests demonstrate that the AWS Clean Rooms collaboration scheme successfully exposed the models to a more complex and realistic variety of advanced threats. The Autoencoder model showed the most robust adaptation in handling combined data dynamics by achieving a prediction precision score of 0.8811 on the 4-Gram sequence extraction. Meanwhile, conventional algorithms experienced a significant performance decline in the collaborative environment due to high false positive rates. This performance degradation was triggered by the inability of conventional models to distinguish lengthy static web asset URLs from actual cyberattack patterns. On the other hand, the Isolation Forest algorithm suffered a detection failure by recording the lowest accuracy scores across almost all scenarios. Therefore, the deployment of the Autoencoder architecture with 4-Gram feature extraction within the AWS Clean Rooms scheme is concluded as the most highly recommended configuration.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Autoencoder, AWS Clean Rooms, Deteksi Anomali, Log Akses Web, Preprocessing Sekuensial, Unsupervised Learning, Autoencoder, AWS Clean Rooms, Anomaly Detection, Web Access Logs, Sequential Preprocessing, Unsupervised Learning |
| Subjects: | Q Science > QA Mathematics > QA336 Artificial Intelligence Q Science > QA Mathematics > QA76.585 Cloud computing. Mobile computing. Q Science > QA Mathematics > QA76.9.A25 Computer security. Digital forensic. Data encryption (Computer science) |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Revy Pramana |
| Date Deposited: | 23 Jul 2026 06:08 |
| Last Modified: | 23 Jul 2026 06:08 |
| URI: | http://repository.its.ac.id/id/eprint/136412 |
Actions (login required)
![]() |
View Item |
