Renggana, Christiant Dimas (2026) Klasifikasi Multi-Label SDGs Pada Artikel Ilmiah Dengan Ekstraksi Kata Kunci Dan Semantic Similarity. Masters thesis, Institut Teknologi Sepuluh Nopember.
|
Text
6025232016-Master_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (1MB) | Request a copy |
Abstract
Universitas memiliki peran penting dalam mendukung pencapaian Sustainable Development Goals (SDGs), terutama melalui riset yang menghasilkan artikel ilmiah. Namun, banyak artikel yang sebenarnya berkontribusi terhadap SDGs tidak teridentifikasi secara sistematis sehingga pemetaan kontribusi universitas menjadi sulit dan kurang akurat. Beberapa penelitian terdahulu yang membuat model untuk klasifikasi SDGs pada artikel ilmiah masih memiliki berbagai research gap, seperti hanya mampu mengklasifikasikan satu label SDGs, hasil klasifikasi yang buruk dengan model teks konvensional, serta penerapan Large Language Model (LLM) yang memerlukan sumber daya komputasi besar dan tidak transparan. Oleh karena itu, penelitian ini bertujuan mengembangkan metode klasifikasi multi-label SDGs berdasarkan judul dan abstrak artikel ilmiah yang akurat sekaligus dapat dipertanggungjawabkan dengan mengintegrasikan ekstraksi kata kunci berbasis semantik dan model transformer.
Dataset terdiri atas 5.765 artikel ilmiah Universitas Airlangga dari platform SciVal yang mencakup SDG 1 hingga SDG 16. Kamus kata kunci SDGs disusun dari query Elsevier. Setiap kata kunci dan setiap artikel kemudian direpresentasikan secara semantik menggunakan Sentence-BERT (SBERT) dengan teknik mean-pooling. Untuk setiap artikel, dipilih sepuluh kata kunci dengan kemiripan tertinggi. Kesepuluh kata kunci tersebut dibandingkan dengan 16 vektor SDGs untuk menghasilkan vektor kemiripan berdimensi 16. Vektor ini berfungsi sebagai auxiliary feature sekaligus dasar penyaringan label awal melalui kalibrasi threshold per-SDG. Selanjutnya, model SciBERT yang telah di-fine-tune diterapkan sebagai multi-label classifier dengan input gabungan judul, abstrak, dan auxiliary feature. Classifier ini dilatih menggunakan multilabel-stratified 5-fold cross validation dan fungsi loss weighted binary cross-entropy untuk mengatasi ketidakseimbangan label. Metode yang diusulkan mencapai macro-F1 0,746 dengan precision 0,748 dan recall 0,751. Pada korpus dan metode training yang sama, metode ini mengungguli baseline TF-IDF dengan Logistic Regression sebesar 11,5 poin dan Word2Vec dengan Logistic Regression sebesar 26,9 poin. Sebagai pembanding, model LLM zero-shot Zephyr 7B mencapai macro-F1 0,717 pada korpus yang berbeda sehingga tidak dapat dibandingkan secara langsung. Meskipun demikian, model tersebut menggunakan parameter sekitar 64 kali lebih banyak dan tidak dapat dijustifikasi. Analisis komponen menunjukkan bahwa fitur kemiripan kata kunci meningkatkan recall secara signifikan, sementara kalibrasi threshold per-SDG memulihkan keseimbangan precision. Dikarenakan setiap prediksi menghasilkan daftar kata kunci SDGs yang cocok, sistem ini mampu menjawab research gap terkait transparansi pada model berbasis LLM.
=============================================================================================================================================
Universities play an important role in supporting the achievement of the Sustainable Development Goals (SDGs), particularly through research that produces scientific articles. However, many articles that contribute to the SDGs are not systematically identified, making it difficult to map university contributions accurately. Previous studies on SDG classification for scientific articles still present several research gaps, including the ability to classify only a single SDG label, poor performance of conventional text models, and the use of Large Language Models (LLMs), which require substantial computational resources and lack transparency. Therefore, this study aims to develop an accurate and accountable multi-label SDG classification method based on article titles and abstracts by integrating semantic keyword extraction and a transformer model.
The dataset consists of 5,765 scientific articles from Universitas Airlangga obtained from SciVal, covering SDG 1 to SDG 16. The SDG keyword dictionary was constructed from Elsevier queries. Each keyword and article was semantically represented using Sentence-BERT (SBERT) with mean pooling. For each article, the ten most similar keywords were selected and compared with 16 SDG vectors to produce a 16-dimensional similarity vector. This vector serves as an auxiliary feature and as the basis for initial label filtering through per-SDG threshold calibration. A fine-tuned SciBERT model was then applied as a multi-label classifier using the combined input of title, abstract, and auxiliary features. The classifier was trained using multilabel-stratified 5-fold cross-validation and weighted binary cross-entropy loss to address label imbalance.
The proposed method achieved a macro-F1 score of 0.746, with precision of 0.748 and recall of 0.751. Under the same corpus and training setting, it outperformed TF-IDF with Logistic Regression by 11.5 points and Word2Vec with Logistic Regression by 26.9 points. As a reference, the zero-shot Zephyr 7B LLM achieved a macro-F1 score of 0.717 on a different corpus, so it cannot be compared directly. Nevertheless, it uses approximately 64 times more parameters and lacks explainability. Component analysis shows that keyword similarity features substantially improve recall, while per-SDG threshold calibration restores precision balance. Since each prediction provides matched SDG keywords, the system addresses the transparency gap in LLM-based models.
| Item Type: | Thesis (Masters) |
|---|---|
| Uncontrolled Keywords: | ekstraksi kata kunci, klasifikasi multi-label, semantic similarity, Sustainable Development Goals, keyword extraction, multi-label classification, semantic similarity |
| Subjects: | T Technology > T Technology (General) > T57.8 Nonlinear programming. Support vector machine. Wavelets. Hidden Markov models. |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Informatics Engineering > 55101-(S2) Master Thesis |
| Depositing User: | Christiant Dimas Renggana |
| Date Deposited: | 30 Jul 2026 04:49 |
| Last Modified: | 30 Jul 2026 04:49 |
| URI: | http://repository.its.ac.id/id/eprint/139550 |
Actions (login required)
![]() |
View Item |
