Krismata, Samuel Yuma (2026) Pencarian Top-K Produk Berbasis Preferensi Pengguna Menggunakan Dynamic Skyline Pada Dataset Terdistribusi. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5027221029-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (5MB) | Request a copy |
Abstract
Pertumbuhan volume data pada berbagai bidang menyebabkan proses pengambilan keputusan multikriteria menjadi semakin kompleks karena setiap objek memiliki kombinasi atribut yang berbeda. Kueri dynamic skyline merupakan salah satu pendekatan yang mampu menghasilkan rekomendasi berdasarkan preferensi pengguna, namun implementasinya pada lingkungan basis data terdistribusi masih menghadapi tantangan dari sisi kompleksitas komputasi. Penelitian ini bertujuan mengimplementasikan pencarian top-k produk berbasis preferensi pengguna menggunakan kueri dynamic skyline pada dataset terdistribusi melalui dua pendekatan komputasi, yaitu SFS-S (Sequential) dan SFS-P (Parallel). Pada kedua pendekatan, komputasi local skyline dilakukan menggunakan algoritma SFS (Sort-Filter-Skyline), sedangkan tahap top-k skyline diimplementasikan menggunakan pendekatan komputasi yang berbeda. Implementasi dilakukan menggunakan Apache Spark, Dask, dan Ray dengan ClickHouse sebagai basis data terdistribusi. Hasil pengujian menunjukkan bahwa kedua metode mampu menghasilkan rekomendasi dengan tingkat akurasi yang tinggi. Berdasarkan rata-rata seluruh skenario pengujian, SFS-P menghasilkan peningkatan nilai precision dan Kendall Tau sebesar 0,03 dibandingkan SFS-S. Dari sisi performa komputasi, SFS-S menunjukkan waktu komputasi yang lebih rendah pada sebagian besar skenario, sedangkan kedua metode mampu menurunkan waktu komputasi hingga 47% dibandingkan algoritma Naïve pada dataset anti-correlated.
=====================================================================================================================================
The growth in data volume across various fields has made multicriteria decision-making increasingly complex, as each object has a different combination of attributes. The dynamic skyline query is one approach capable of generating recommendations based on user preferences; however, its implementation in a distributed database environment still faces challenges in terms of computational complexity. This research aims to implement a top-k product search based on user preferences using dynamic skyline queries on a distributed dataset through two computational approaches, namely SFS-S (Sequential) and SFS-P (Parallel). In both approaches, local skyline computation is performed using the SFS (Sort-Filter-Skyline) algorithm, while the top-k skyline stage is implemented using different computational approaches. The implementation was carried out using Apache Spark, Dask, and Ray with ClickHouse as the distributed database. Test results show that both methods are capable of generating recommendations with a high level of accuracy. Based on the average across all test scenarios, SFS-P yields an increase in precision and Kendall’s tau of 0.03 compared to SFS-S. In terms of computational performance, SFS-S achieved shorter computation times in most scenarios, while both methods were able to reduce computation time by up to 47% compared to the Naïve algorithm on the anti-correlated dataset.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Apache Spark, Basis Data Terdistribusi, Dask, Dynamic Skyline, Ray, Distributed Database |
| Subjects: | T Technology > T Technology (General) > T57.5 Data Processing |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information Technology > 59201-(S1) Undergraduate Thesis |
| Depositing User: | Samuel Yuma Krismata |
| Date Deposited: | 20 Jul 2026 03:31 |
| Last Modified: | 21 Jul 2026 02:00 |
| URI: | http://repository.its.ac.id/id/eprint/135548 |
Actions (login required)
![]() |
View Item |
