Putra, Muhammad Farhan Lucky (2026) Analisis Pemodelan Topik Ulasan Aplikasi Sekuritas Menggunakan Latent Dirichlet Allocation Dan Non-negative Matrix Factorization. Diploma thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5003221001-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (11MB) | Request a copy |
Abstract
Pertumbuhan investor pasar modal Indonesia yang menembus 20 juta SID pada tahun 2025 mendorong peningkatan penggunaan aplikasi sekuritas berbasis seluler, khususnya Stockbit dan Indo Premier Sekuritas (IPOT). Ulasan pengguna yang terkumpul di Google Play Store mengandung opini terhadap berbagai aspek layanan secara bersamaan dan bersifat teks tidak terstruktur, sehingga diperlukan pendekatan otomatis untuk mengidentifikasi topik-topik utama yang dibahas pengguna secara sistematis. Penelitian ini menerapkan pemodelan topik menggunakan Latent Dirichlet Allocation (LDA) dan Non-Negative Matrix Factorization (NMF) dengan dua representasi teks yaitu Bag of Words (BoW) dan TF-IDF terhadap 10.676 ulasan berbahasa Indonesia periode 2025–2026 yang diperoleh melalui web scraping dari Google Play Store. Tahap preprocessing menghasilkan 6.274 kalimat valid yang digunakan sebagai input pemodelan. Keempat kombinasi model diuji dengan jumlah topik optimal K = 4 yang ditentukan berdasarkan grid search menggunakan coherence score dan topic diversity. LDA BoW dilatih dengan hyperparamer α = 0,15 dan β = 0,01, sedangkan LDA TF-IDF dengan α = 0,15 dan β = 0,01. Keempat kombinasi model secara konsisten mengidentifikasi empat topik utama yaitu Pengalaman Pengguna, Registrasi & Layanan Pelanggan, Transaksi & Performa Aplikasi, dan Fitur & Edukasi Investasi. NMF TF-IDF menghasilkan coherence score tertinggi sebesar 0,7133 dan topic diversity sempurna sebesar 1,0000, diikuti LDA TF-IDF dengan coherence score 0,6700 dan topic diversity 0,9000, NMF BoW dengan coherence score 0,6682 dan topic diversity 0,9500, dan LDA BoW dengan coherence score 0,5980 dan topic diversity 0,8000. Keunggulan NMF TF-IDF konsisten dengan karakteristiknya yang mampu menghasilkan pemisahan topik yang lebih distinct pada data teks pendek hasil sentence segmentation serta keunggulan representasi TF-IDF dalam menekan kata-kata umum yang tidak diskriminatif. Temuan ini memberikan wawasan berbasis data bagi pengembang Stockbit dan IPOT dalam mengidentifikasi aspek-aspek layanan yang menjadi perhatian utama pengguna sebagai dasar evaluasi dan peningkatan kualitas layanan.
====================================================================================================================================
The growth of Indonesian capital market investors surpassing 20 million SID in 2025 has driven increased usage of mobile-based securities applications, particularly Stockbit and Indo Premier Sekuritas (IPOT). User reviews collected on Google Play Store contain opinions on various aspects of services simultaneously and are unstructured text data, requiring an automated approach to systematically identify the main topics discussed by users. This study applies topic modeling using Latent Dirichlet Allocation (LDA) and Non-Negative Matrix Factorization (NMF) with two text representations, namely Bag of Words (BoW) and TF-IDF, on 10,676 Indonesian-language reviews from the period 2025–2026 obtained through web scraping from Google Play Store. The preprocessing stage produced 6,274 valid sentences used as modeling input. All four model combinations were tested with the optimal number of topics K = 4, determined through grid search using coherence score and topic diversity. LDA BoW was trained with hyperparameters α = 0.15 and β = 0.01, while LDA TF-IDF with α = 0.05 and β = 0.05. All four model combinations consistently identified four main topics: User Experience, Registration & Customer Service, Transaction & Application Performance, and Features & Investment Education. NMF TF-IDF produced the highest coherence score of 0.7133 and a perfect topic diversity of 1.0000, followed by LDA TF-IDF with coherence score 0.6700 and topic diversity 0.9000, NMF BoW with coherence score 0.6682 and topic diversity 0.9500, and LDA BoW with coherence score 0.5980 and topic diversity 0.8000. The superiority of NMF TF-IDF is consistent with its ability to produce more distinct topic separation on short text data resulting from sentence segmentation, as well as the advantage of TF-IDF representation in suppressing non-discriminative common words. These findings provide data-driven insights for Stockbit and IPOT developers in identifying key service aspects of concern to users as a basis for service quality evaluation and improvement.
| Item Type: | Thesis (Diploma) |
|---|---|
| Uncontrolled Keywords: | Latent Dirichlet Allocation, Non-Negative Matrix Factorization, Pemodelan Topik, Sentence Segmentation, Ulasan Aplikasi Sekuritas, Latent Dirichlet Allocation, Non-Negative Matrix Factorization, Securities Application Reviews, Sentence Segmentation, Topic Modeling. |
| Subjects: | Q Science > Q Science (General) > Q325.5 Machine learning. Support vector machines. Q Science > QA Mathematics > QA278.55 Cluster analysis Q Science > QA Mathematics > QA76.9.D343 Data mining. Querying (Computer science) |
| Divisions: | Faculty of Science and Data Analytics (SCIENTICS) > Statistics > 49201-(S1) Undergraduate Thesis |
| Depositing User: | Muhammad Farhan Lucky Putra |
| Date Deposited: | 05 Aug 2026 03:39 |
| Last Modified: | 05 Aug 2026 03:39 |
| URI: | http://repository.its.ac.id/id/eprint/143851 |
Actions (login required)
![]() |
View Item |
