Clustering Berdasarkan Graph Database untuk Pencarian Data Pada Pipeline Repository

Alfian, Moh Adam (2019) Clustering Berdasarkan Graph Database untuk Pencarian Data Pada Pipeline Repository. Other thesis, Institut Teknologi Sepuluh Nopember.

[thumbnail of 05111540007005-Undergraduate_Thesis.pdf] Text
05111540007005-Undergraduate_Thesis.pdf
Restricted to Repository staff only

Download (3MB) | Request a copy

Abstract

Kondisi saat ini pencarian data pada Pipeline Repository hanya mampu mencari pipeline berdasarkan judul pipeline. Pencarian berdasarkan judul pipeline mempunyai dua kelemahan, yang pertama pemberian judul pipeline terkadang tidak merepresentasikan isi pipeline, sehingga akan menyulitkan pengguna dalam menemukan desain pipeline yang diinginkan. Kelemahan kedua yaitu pencarian berdasarkan judul belum mampu menangani masalah ketika pengguna hanya mengingat proses-proses yang terjadi di dalamnya dan tidak mengetahui judul pipeline untuk proses tersebut. Untuk itu dalam Tugas Akhir ini akan dibuat pencarian pipeline berdasarkan stage yang terkandung pada pipeline menggunakan basis data Neo4J dan MySQL. Cypher dan GraphQL digunakan sebagai bahasa query. Selain itu dalam Tugas Akhir ini juga akan membandingkan kemiripan desain pipeline dengan pipeline yang lain menggunakan algoritme Jaccard Similarity dan dilakukan pengelompokan menggunakan algoritme K-medoids dan menghitung biaya edit tree untuk mengubah pipeline menjadi pipeline yang lain menggunakan algoritme Tree Edit Distance. Dengan adanya aplikasi ini, pengguna dapat mencari stage apa saja yang terkandung pada setiap pipeline, mengetahui tingkat kemiripan tiap pipeline, serta mengetahui biaya edit untuk mengubah suatu pipeline menjadi pipeline yang lain menggunakan algoritma Tree Edit Distance.
===================================================================================================================================
The current condition of data search on the Pipeline Repository is only able to find a pipeline based on the pipeline title. Search based on the pipeline title has two disadvantages, the first is the title of the pipeline sometimes does not represent the contents of the pipeline, so it will be difficult for users to find the desired pipeline design. The second weakness is the search based on the title has not been able to handle the problem when the user only remembers the processes that occur in it and does not know the title of the pipeline for the process. For this reason, in this final project a pipeline search will be made based on the stages contained in the pipeline using Neo4J and MySQL databases. Cypher and GraphQL will be used for query languages. In addition, this final project will also compare the similarity of the pipeline design with the other pipelines using the Jaccard Similarity algorithm, group them using the K-medoids algorithm and calculate the cost of editing trees to change a pipeline to another pipeline using the Tree Edit Distance algorithm. With this application, users can search for any stage contained in each pipeline, know the level of similarity of each pipeline, and find out the editing costs to change a pipeline to another pipeline using the Tree Edit Distance algorithm.

Item Type: Thesis (Other)
Uncontrolled Keywords: Jaccard Similarity, K-Medoids, Pipelines, Streamsets Data Collector, Tree Edit Distance
Subjects: T Technology > T Technology (General) > T58.6 Management information systems
Divisions: Faculty of Information and Communication Technology > Informatics > 55201-(S1) Undergraduate Thesis
Depositing User: Moh Adam Alfian
Date Deposited: 23 Jul 2026 08:08
Last Modified: 23 Jul 2026 08:08
URI: http://repository.its.ac.id/id/eprint/70367

Actions (login required)

View Item View Item