Arrasyid, Vidiawan Nabiel (2026) Analisis Kinerja Large Language Model dalam Tugas Multi-Document Summarization dengan Pendekatan Abstraktif dan Ekstraktif. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5025221231-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (8MB) | Request a copy |
Abstract
Pertumbuhan data teks yang pesat memunculkan permasalahan information overload, dan multi-document summarization menjadi solusi untuk menyajikan informasi secara ringkas dari sekumpulan dokumen redundan. Large Language Model (LLM) menjanjikan pada tugas ini melalui pendekatan abstraktif maupun ekstraktif, namun perbandingan sistematis kedua pendekatan pada model yang sama masih terbatas dan pengaruh panjang konteks belum banyak diperiksa secara terkendali. Penelitian ini membandingkan pendekatan abstraktif dan ekstraktif pada lima open-weight LLM (gemma-3-4b, llama-3.1-8b, mistral-nemo, phi-4-mini, qwen3-8b) secara zero-shot pada dataset Multi-News, dievaluasi melalui kerangka empat aspek: leksikal (ROUGE, METEOR), semantik (BERTScore), faktualitas (MiniCheck), dan karakteristik tekstual. Penelitian dilakukan dalam tiga tahapan: studi pendahuluan merancang dan memvalidasi prompt zero-shot untuk kedua pendekatan, Eksperimen 1 membandingkan kedua pendekatan pada seluruh model dan menetapkan model terbaik, sedangkan Eksperimen 2 menganalisis pengaruh jumlah dokumen per klaster berita pada kombinasi model dan pendekatan terbaik, dengan uji nonparametrik, ukuran efek, dan koreksi perbandingan ganda. Kedua pendekatan unggul pada aspek yang berbeda. Pendekatan ekstraktif lebih unggul pada aspek faktualitas serta sebagian besar metrik leksikal, terutama ROUGE-2 dan METEOR, secara konsisten di kelima model dan unggul pada 59,4 persen kombinasi metrik dan model, sedangkan abstraktif lebih unggul pada aspek semantik, dan llama-3.1-8b terpilih sebagai model terbaik melalui kriteria faktualitas dan pemutus seri multi-aspek. Kualitas ringkasan ekstraktif menurun konsisten seiring bertambahnya jumlah dokumen, paling tajam pada aspek semantik, diikuti faktualitas yang turun dari 0,838 menjadi 0,665.
==================================================================================================================================
The rapid growth of textual data has led to the problem of information overload, and multi-document summarization has become a solution for presenting concise information from a set of redundant documents. Large Language Models (LLMs) are promising for this task through both abstractive and extractive approaches, yet systematic comparisons of the two approaches on the same models remain limited and the effect of context length has rarely been examined in a controlled manner. This study compares the abstractive and extractive approaches on five open-weight LLMs (gemma-3-4b, llama-3.1-8b, mistral-nemo, phi-4-mini, qwen3-8b) in a zero-shot setting on the Multi-News dataset, evaluated through a four-aspect framework: lexical (ROUGE, METEOR), semantic (BERTScore), factuality (MiniCheck), and textual characteristics. The study is conducted in three stages: a preliminary study designs and validates the zero-shot prompts for both approaches, Experiment 1 compares the two approaches across all models and selects the best model, while Experiment 2 analyzes the effect of the number of documents per news cluster on the best model-approach combination, using nonparametric tests, effect sizes, and multiple-comparison corrections. The two approaches excel on different aspects. The extractive approach is superior on the factuality aspect and on most lexical metrics, particularly ROUGE-2 and METEOR, consistently across all five models, and is superior on 59.4 percent of metric-model combinations, while the abstractive approach is superior on the semantic aspect, and llama-3.1-8b is selected as the best model through a factuality criterion and a multi-aspect tie-breaker. The quality of extractive summaries declines consistently as the number of documents increases, most sharply on the semantic aspect, followed by factuality, which drops from 0.838 to 0.665.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Large Language Model, Multi-Document Summarization, Abstraktif, Ekstraktif, Faktualitas. Large Language Model, Multi-Document Summarization, Abstractive, Extractive, Factuality. |
| Subjects: | Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. Q Science > QA Mathematics > QA336 Artificial Intelligence Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science) |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Vidiawan Nabiel Arrasyid |
| Date Deposited: | 31 Jul 2026 03:11 |
| Last Modified: | 31 Jul 2026 03:11 |
| URI: | http://repository.its.ac.id/id/eprint/138605 |
Actions (login required)
![]() |
View Item |
