Penerapan Algoritma Genetika Dalam Pemeringatan TERM pada Klasterisasi Hasil Pencarian WEB

Tua, Maruli (2009) Penerapan Algoritma Genetika Dalam Pemeringatan TERM pada Klasterisasi Hasil Pencarian WEB. Masters thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 5104201020-Master_thesis.pdf] Text
5104201020-Master_thesis.pdf
Restricted to Repository staff only

Download (17MB)

Abstract

Pemeringkatan term merupakan salah satu proses penting yang dapat membantu menghasilkan klaster dokumen dengan kualitas yang baik. Ketika klasterisasi dilakukan pada sekumpulan dokumen web dengan ukuran yang sangat besar dan jumlah term yang juga banyak, kinerja dan kualitas proses klasterisasi dapat menurun. Berbagai pengembangan metode yang ditujukan untuk mengatasi persoalan ini merupakan salah satu ranah penelitian yang menarik dalam bidang temu kembali informasi. Dalam penelitian ini, algoritma pemeringkatan term berbasis algoritma genetika (GA) dikembangkan dengan tujuan untuk meningkatkan kualitas proses klasterisasi sekumpulan dokumen web yang didasarkan pada kata-kata (term) kunci yang terkandung dalam dokumen yang akan diklaster. Dalam algoritma yang dikembangkan, semua term hasil ekstraksi dari dokumen web direpresentasikan dalam bentuk relasi graf tidak berarah (relational undirected graph), di mana setiap simpul dan busur dari graf berturut-turut merepresentasikan term dan hubungan (relasi) antarterm. Algoritma genetika digunakan untuk memperoleh graf yang berisikan term dan hubungan antarterm yang optimal. Optimalitas graf yang terbentuk diukur berdasarkan dua faktor utama yang berpengaruh dalam pemeringkatan term, yaitu important neighbors dan strong connections. Algoritma genetika berperan dalam menentukan bobot simpul dan busur dari graf. Fungsi fitness dari GA yang digunakan didefinisikan sebagai fungsi beberapa parameter yang dapat digunakan untuk mengukur kualitas klaster dokumen. Algoritma pemeringkatan term berbasis GA yang dikembangkan dalam penelitian ini diuji coba dengan beberapa dokumen yang diperoleh dari Internet. Nilai parameter genetika terbaik yang dihasilkan dalam uji coba digunakan untuk mengukur kualitas klasterisasi pencarian web. Hasil uji coba perbandingan menunjukkan bahwa penggunaan GA mampu meningkatkan kualitas klasterisasi dibandingkan dengan hasil yang diperoleh menggunakan metode TF-IDF. Untuk itu, nilai precision dan recall rata-rata sebesar 80,7% dan 26,6% diperoleh untuk GA, sedangkan nilai rata-rata precision dan recall sebesar 72,6% dan 21,3% diperoleh untuk TF-IDF.
==================================================================================================================================
Terms ranking is one of the important processes that can help produce good-quality document clusters. When clustering is performed on a very large set of web documents containing a large number of terms, the performance and quality of the clustering process may decrease. Therefore, the development of various methods intended to overcome this problem has become an interesting research area in the field of information retrieval. In this research, a term-ranking algorithm based on the genetic algorithm (GA) was developed with the aim of improving the quality of the clustering process for a set of web documents based on the key terms contained in the documents. In the developed algorithm, all terms extracted from web documents are represented as a relational undirected graph, where each node and edge of the graph corresponds to a term and a relationship between a pair of terms, respectively. A genetic algorithm is employed to produce an optimal relational undirected graph. The optimality of the graph is measured based on two main factors that influence the term-ranking process, namely important neighbors and strong connections. The genetic algorithm plays a role in determining the weights of the nodes and edges of the graph. The fitness function of the GA is defined as a function of several parameters that can be used to measure the quality of document clusters. The GA-based term-ranking algorithm developed in this research was tested using several web documents obtained from the Internet. The best genetic parameter values obtained from the tests were used to measure the quality of web search clustering. The comparison test results showed that the use of GA was able to improve clustering quality compared with the results obtained using the TF-IDF method. In this regard, average precision and recall values of 80.7% and 26.6%, respectively, were obtained for the GA-based algorithm, while average precision and recall values of 72.6% and 21.3%, respectively, were obtained for the TF-IDF method.

Item Type: Thesis (Masters)
Additional Information: RSE 621.384 Sib e
Uncontrolled Keywords: klasterisasi dokumen, pemeringkatan term, algoritma genetika, relasi graftidak berarah, document clustering, terms ranking, genetic algorithms, relationalundirected graph.
Subjects: T Technology > TK Electrical engineering. Electronics Nuclear engineering > TK5103.2 Wireless communication systems. Two way wireless communication
Divisions: Faculty of Information Technology > Informatics Engineering > 55101-(S2) Master Thesis
Depositing User: magang .
Date Deposited: 18 Sep 2026 02:46
Last Modified: 18 Sep 2026 02:49
URI: http://repository.its.ac.id/id/eprint/144663

Actions (login required)

View Item View Item