Purwito, Luthfan Aryananda (2026) Analisis Kinerja Arsitektur Integrasi Data ETL, ELT, dan Hibrida pada Data Warehouse Berbasis PostgreSQL. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5026221166-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (2MB) | Request a copy |
Abstract
Pertumbuhan volume dan varietas data menuntut organisasi untuk memiliki strategi integrasi data yang efisien. Arsitektur Extract, Transform, Load (ETL) telah lama menjadi standar dalam integrasi data, namun menghadapi keterbatasan dari sisi skalabilitas dan kinerja ketika menangani data bervolume besar. Sebagai alternatif, arsitektur Extract, Load, Transform (ELT) memanfaatkan kekuatan komputasi mesin basis data target, sementara arsitektur hibrida berupaya menggabungkan keunggulan keduanya. Namun, studi komparatif yang mengukur kinerja ketiga arsitektur ini secara kuantitatif pada lingkungan lokal (on-premise) masih terbatas. Penelitian ini bertujuan untuk mengevaluasi dan membandingkan kinerja arsitektur ETL, ELT, dan hibrida dalam mengolah data terstruktur (CSV dan Parquet) serta semi-terstruktur (JSON) pada data warehouse berbasis PostgreSQL. Arsitektur ETL diimplementasikan menggunakan Python (PySpark), sedangkan ELT dan hibrida menggunakan kombinasi Python dan dbt (data build tool). Pengujian dilakukan menggunakan dataset benchmark TPC-H pada dua skala volume (SF 5 dan SF 10) dengan metrik waktu pemrosesan, efisiensi sumber daya komputasi, penggunaan ruang penyimpanan, serta skalabilitas. Hasil penelitian menunjukkan bahwa arsitektur ETL secara konsisten memiliki waktu pemrosesan tercepat dan paling tahan terhadap peningkatan volume data, diikuti arsitektur hibrida, sedangkan ELT paling lambat. Titik hambatan (bottleneck) utama pada ELT dan hibrida terletak pada tahap pemuatan data ke basis data melalui JDBC, bukan pada tahap transformasi. Dari sisi penyimpanan, format Parquet terbukti paling efisien sedangkan JSON paling boros, dan arsitektur ELT meninggalkan jejak tabel mentah terbesar. Temuan ini diharapkan dapat menjadi referensi dalam pemilihan arsitektur integrasi data yang sesuai dengan karakteristik data dan lingkungan pemrosesan.
=======================================================================================================================================
The growth in data volume and variety requires organizations to adopt efficient data integration strategies. The Extract, Transform, Load (ETL) architecture has long been the standard for data integration; however, it faces limitations in scalability and performance when handling large-volume data. As an alternative, the Extract, Load, Transform (ELT) architecture leverages the computational power of the target database engine, while the hybrid architecture attempts to combine the strengths of both. Nevertheless, comparative studies that quantitatively measure the performance of these three architectures in on-premise environments remain limited. This study aims to evaluate and compare the performance of the ETL, ELT, and hybrid architectures in processing structured (CSV and Parquet) and semi-structured (JSON) data within a PostgreSQL-based data warehouse. The ETL architecture is implemented using Python (PySpark), while the ELT and hybrid architectures use a combination of Python and dbt (data build tool). The evaluation was conducted using the TPC-H benchmark dataset at two data scales (SF 5 and SF 10), measuring processing time, computational resource efficiency, storage usage, and scalability. The results show that the ETL architecture consistently achieves the fastest processing time and is the most resilient to increasing data volume, followed by the hybrid architecture, while ELT is the slowest. The main bottleneck in the ELT and hybrid architectures lies in the data loading stage to the database via JDBC, rather than in the transformation stage. In terms of storage, the Parquet format proves to be the most efficient while JSON is the most costly, and the ELT architecture leaves the largest raw table footprint. These findings are expected to serve as a reference for selecting a data integration architecture suited to the characteristics of the data and the processing environment.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Integrasi Data, ETL, ELT , Hibrida, PostgreSQL, dbt, Python, TPC-H, Data Integration, ETL, ELT , Hybrid, PostgreSQL, dbt, Python, TPC-H |
| Subjects: | Q Science > QA Mathematics > QA76.9.D37 Data warehousing. Q Science > QA Mathematics > QA76.9D338 Data integration |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information System > 57201-(S1) Undergraduate Thesis |
| Depositing User: | Luthfan Aryananda Purwito |
| Date Deposited: | 29 Jul 2026 01:11 |
| Last Modified: | 29 Jul 2026 01:11 |
| URI: | http://repository.its.ac.id/id/eprint/139374 |
Actions (login required)
![]() |
View Item |
