Prabowo, Angela Oryza (2026) Evaluasi Kemampuan Large Language Model untuk Penalaran Lintas Sumber Cyber Threat Intelligence dan Deteksi Dini Emerging Campaign. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6025241038-Master_Thesis.pdf Restricted to Repository staff only Download (2MB) | Request a copy |
Abstract
Laporan CTI merupakan sumber informasi penting dalam memahami ancaman siber. Namun, pada tahap awal kemunculan sebuah emerging campaign, informasi mengenai campaign tersebut umumnya masih tersebar pada berbagai laporan CTI dari vendor keamanan siber yang berbeda, menggunakan terminologi yang tidak seragam, dan hanya menggambarkan sebagian dari keseluruhan aktivitas serangan. Kondisi ini menyulitkan analis untuk mengidentifikasi keterkaitan antar laporan serta membangun pemahaman yang komprehensif mengenai campaign yang sedang berkembang. Meskipun Large Language Models (LLMs) menunjukkan kemampuan yang menjanjikan dalam berbagai tugas pemrosesan bahasa alami, belum terdapat kerangka evaluasi yang secara khusus mengukur kemampuan LLM dalam memahami emerging campaign melalui penalaran lintas laporan. Penelitian ini mengusulkan kerangka evaluasi dua tahap untuk mengukur kemampuan tersebut, yaitu pengelompokan emerging campaign, yang mengevaluasi kemampuan LLM mengidentifikasi laporan CTI yang membahas emerging campaign yang sama, serta sintesis fungsional, yang mengevaluasi kemampuan LLM mengintegrasikan informasi dari berbagai laporan menjadi sebuah laporan CTI yang komprehensif. Untuk mendukung evaluasi, penelitian ini membangun dataset yang terdiri atas 375 laporan CTI, mencakup 175 laporan emerging campaign yang dikelompokkan ke dalam 10 emerging campaign serta 200 laporan noise. Evaluasi dilakukan terhadap tujuh LLM menggunakan metrik F1 pada tahap pengelompokan emerging campaign dan coverage berbasis MS-MARCO pada tahap sintesis fungsional. Hasil penelitian menunjukkan bahwa pada kelompok LLM proprietary, NotebookLM memperoleh rata-rata nilai F1 tertinggi sebesar 0,951 pada tahap pengelompokan emerging campaign, sedangkan Gemini memperoleh rata-rata coverage tertinggi sebesar 55,6% pada tahap sintesis fungsional. Pada kelompok LLM lokal, Llama memperoleh rata-rata nilai F1 tertinggi sebesar 0,316, sedangkan Qwen memperoleh rata-rata coverage tertinggi sebesar 19,4%. Analisis lebih lanjut menunjukkan bahwa keberhasilan deteksi dipengaruhi oleh karakteristik campaign, sementara prior knowledge model tidak memiliki hubungan yang signifikan secara statistik terhadap performa pengelompokan emerging campaign.
======================================================================================================================================
Cyber Threat Intelligence (CTI) reports are a primary source of information for understanding cyber threats. However, during the early stages of an emerging campaign, intelligence is typically fragmented across reports published by different security vendors, described using inconsistent terminology, and often captures only partial observations of the campaign. This fragmentation makes it difficult for analysts to identify relationships among reports and construct a comprehensive understanding of the campaign. Although Large Language Models (LLMs) have demonstrated strong capabilities in various natural language processing tasks, there is currently no dedicated evaluation framework for assessing their ability to understand emerging campaigns through cross-document reasoning. This study proposes a two-stage evaluation framework to measure this capability. The first stage, Emerging Campaign Grouping, evaluates an LLM's ability to identify CTI reports describing the same emerging campaign while distinguishing them from unrelated reports. The second stage, Functional Synthesis, evaluates the model's ability to integrate information distributed across multiple CTI reports into a comprehensive campaign report. To support the evaluation, we constructed a dataset consisting of 375 CTI reports, including 175 reports covering 10 emerging campaigns and 200 carefully curated noise reports. Seven LLMs, including both proprietary and self-hosted models, were evaluated using the F1-score for Emerging Campaign Grouping and an MS-MARCO-based coverage metric for Functional Synthesis. The experimental results show that among the proprietary LLMs, NotebookLM achieved the highest average F1-score of 0.951 in the emerging campaign grouping task, while Gemini achieved the highest average coverage of 55.6% in the functional synthesis task. Among the local LLMs, Llama obtained the highest average F1-score (0.316), whereas Qwen achieved the highest average coverage (19.4%). Further analysis indicates that detection performance is influenced by the characteristics of the campaign, whereas the models' prior knowledge does not exhibit a statistically significant relationship with emerging campaign grouping performance.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | Cyber Threat Intelligence, Deteksi Dini Kampanye Siber, Large Language Model, Penalaran Lintas Laporan, Cyber Threat Intelligence, Cross-Report Reasoning, Early Campaign Detection, Large Language Models |
| Subjects: | Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. Q Science > Q Science (General) > Q337.5 Pattern recognition systems T Technology > T Technology (General) > T57.5 Data Processing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis |
| Depositing User: | Angela Oryza Prabowo |
| Date Deposited: | 03 Aug 2026 05:36 |
| Last Modified: | 03 Aug 2026 05:36 |
| URI: | http://repository.its.ac.id/id/eprint/142168 |
Actions (login required)
![]() |
View Item |
