Pembobotan Term Berbasis Fitur Named Entity dan Posisi Term untuk Klasifikasi Multi-label Artikel Berita

Fikri, Mohammad (2019) Pembobotan Term Berbasis Fitur Named Entity dan Posisi Term untuk Klasifikasi Multi-label Artikel Berita. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 05111540000049-Undergraduate_Theses.pdf] Text
05111540000049-Undergraduate_Theses.pdf
Restricted to Repository staff only

Download (4MB) | Request a copy

Abstract

Penentuan kategori sebuah artikel berita merupakan salah satu tugas utama dari sekian banyaknya tugas yang harus dikerjakan oleh seorang redaktur. Seiring dengan berkembangnya zaman, para redaktur berhadapan dengan artikel berita yang sangat bervariasi dan tidak semuanya masuk ke dalam satu kategori saja. Salah satu cara yang dapat dilakukan dalam penentuan kategori adalah melihat term yang dianggap sebagai kata kunci pada konten artikel berita. Kata kunci bisa merupakan sebuah term yang merupakan entitas bernama (named entity) seperti nama orang, nama organisasi, nama lokasi ataupun nama event. Judul artikel berita juga bisa menjadi salah satu penentu kategori dari artikel berita karena judul merupakan representasi dari konten artikel berita. Sehingga diperlukan sebuah cara agar pembobotan fitur yang digunakan lebih mempertimbangkan term yang merupakan named entity ataupun term yang terdapat pada judul artikel berita lebih dipertimbangkan dibandingkan term lainnya dalam proses penentuan kategori dari artikel berita. Pada tugas akhir ini diusulkan sebuah sistem yang dapat mengklasifikasikan sebuah artikel berita ke dalam kategori-kategori relevan dengan menggunakan pembobotan fitur berbasis fitur named entity dan posisi term. Data yang digunakan merupakan artikel berita berbahasa Indonesia dari beberapa portal berita online Indonesia. Untuk dapat melakukan klasifikasi artikel berita digunakan metode deep learning dengan arsitektur XML-CNN. Agar dapat mengakomodasi tujuan utama dari tugas akhir, diusulkan sebuah metode yang merupakan modifikasi dari dua metode pembobotan fitur, yaitu Confidence Weight dan Title Weight. Dari hasil pengujian diketahui bahwa sistem yang diusulkan dapat dengan baik mengklasifikasi artikel berita ke dalam kategori-kategori yang relevan dengan performa recall sebesar 93.70%, precision sebesar 93.72%, f-measure sebesar 93.64%, accuracy sebesar 93.70%, dan nilai loss sebesar 0.03535.
===================================================================================================================================
Determining the category of a news article is one of the main tasks that must be done by an editor. Along with the development of the times, editors deal with news articles that are very varied and not all fall into one category only. One way that can be done in determining categories is to look at terms that are considered as keywords in the content of news articles. Keywords can be a term which is a named entity such as a person's name, organization name, location name or event name. The title of a news article can also be a determinant of a category from a news article because the title is a representation of the content of a news article. Therefore, a method is needed so that a term that is named entity or contained in the title of a news article is more considered than other terms in the process of determining categories from news articles. In this project, a system that can classify a news article into relevant categories by using term weighting based on named entity features and term positions is proposed. The data used is Indonesian language news articles from several Indonesian online news portals. To be able to classify news articles, the deep learning method is used with the XML-CNN architecture. In order to accommodate the main objectives of the final assignment, a method is proposed which is a modification of two feature weighting methods, namely Confidence Weight and Title Weight. From the test results it is known that the proposed system can properly classify news articles into relevant categories with a recall performance of 93.70%, precision of 93.72%, f-measure of 93.64%, accuracy of 93.70%, and loss value of 0.03535.

Item Type: Thesis (Other)
Uncontrolled Keywords: artikel berita, Confidence Weight, klasifikasi multi-label, named entity, Title Weight, XML-CNN
Subjects: Q Science > QA Mathematics > QA336 Artificial Intelligence
Q Science > QA Mathematics > QA76.87 Neural networks (Computer Science)
Q Science > QA Mathematics > QA76.9.D343 Data mining. Querying (Computer science)
T Technology > T Technology (General) > T57.5 Data Processing
Divisions: Faculty of Information and Communication Technology > Informatics > 55201-(S1) Undergraduate Thesis
Depositing User: Mohammad Fikri
Date Deposited: 23 Jul 2026 04:23
Last Modified: 23 Jul 2026 04:23
URI: http://repository.its.ac.id/id/eprint/65808

Actions (login required)

View Item View Item