Klasterisasi Topik pada Artikel Ilmiah menggunakan Weighted K-means dan Maximum Common Subgraph

Muhammad, Riduwan (2019) Klasterisasi Topik pada Artikel Ilmiah menggunakan Weighted K-means dan Maximum Common Subgraph. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5116201019_Master_Thesis.pdf] Text
5116201019_Master_Thesis.pdf
Restricted to Repository staff only

Download (2MB) | Request a copy

Abstract

Jumlah penelitian di dunia mengalami perkembangan yang pesat, setiap tahun berbagai peneliti dari penjuru dunia menghasilkan publikasi ilmiah seperti makalah, jurnal, buku dsb. Metode klasterisasi dapat digunakan untuk mengelompokkan kumpulan dokumen publikasi ilmiah ke dalam suatu kelompok tertentu berdasarkan relevansi antar topik. Klasterisasi pada dokumen memiliki karakteristik yang berbeda karena tingkat kemiripan antar dokumen dipengaruhi oleh kata-kata pembentuknya. Beberapa metode klasterisasi kurang memperhatikan nilai semantik dari kata. Sehingga klaster yang terbentuk kurang merepresentasikan isi topik dokumen. Klasterisasi dokumen teks masih memiliki kemungkinan adanya outlier karena pemilihan fitur teks yang tidak optimal. Oleh karena itu dibutuhkan pemrosesan data yang tepat serta metode yang mengoptimalkan hasil klaster. Penelitian ini mengusulkan metode klasterisasi dokumen menggunakan Weighted K-Means yang dipadukan dengan Maximum Common Subgraph. Weighted K-means digunakan untuk klasterisasi awal dokumen berdasarkan kata-kata yang diekstraksi. Pembentukan Weighted K-Means berdasarkan perhitungan Word2Vec dan TextRank dari kata-kata dalam dokumen. Maximum common subgraph merupakan tahap pembentukan graf yang digunakan dalam penggabungan klaster untuk menghasilkan klaster baru yang lebih optimal. Pembentukan graf dilakukan dengan perhitungan nilai Word2Vec dan Word Co-occurrence dari klaster. Representasi topik dokumen tiap klaster dapat dihasilkan dari pemodelan topik Latent Dirichlet Allocation (LDA). Pengujian dilakukan dengan mengukur Silhouette Coefficient untuk mengukur nilai k optimal. Selanjutnya analisis Koherensi topik membandingkan antara usulan metode dengan algoritma k-means dasar.
===================================================================================================================================
The number of researches in the world has been increased, every year researchers from around the world produce scientific publications such as journals, proceeding, books, etc. Clustering methods can be used to cluster scientific publications based on relevance between topics. Clustering on document has different characteristics because the level of similarity between documents is influenced by the words. Some clustering methods less attention to the semantic value of the word. The cluster formed does not represent the topic of the document. Document clustering still has the possibility of outliers because the selection of text features is not optimal. Therefore, proper data processing and methods that optimize cluster results are needed. This research proposes a document clustering method using Weighted K-Means and Maximum Common Subgraph. Weighted K-means are used for initial clustering of documents based on extracted words. Weighted K-Means formed by Word2Vec and TextRank. Then Maximum Common Subgraph is the graph formation stage used in combining clusters to produce a new optimal cluster. Graph is formed by Word2vec similarity and Word Co-occurrence of clusters. The topic cluster can be generated from the modeling of the Latent Dirichlet Allocation (LDA). Tests carried out by measuring the Silhouette Coefficient to measure the optimal k value. Furthermore, the topic Coherence analysis compares the proposed method with the basic k-means algorithm.

Item Type: Thesis (Masters)
Uncontrolled Keywords: Latent Dirichlet Allocation, Maximum Common Subgraph, TextRank, Weighted K-means, Word Co-occurence, Word2Vec
Subjects: T Technology > T Technology (General) > T57.5 Data Processing
T Technology > T Technology (General) > T58.5 Information technology. IT--Auditing
Divisions: Faculty of Information Technology > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: Muhammad Riduwan
Date Deposited: 20 Jul 2026 01:43
Last Modified: 20 Jul 2026 01:43
URI: http://repository.its.ac.id/id/eprint/70624

Actions (login required)

View Item View Item