Zaqiyah, Ana Alimatus (2019) Deteksi Opini Spam pada review produk Menggunakan Support Vector Machine. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
05111540000115-Undergraduate_Theses.pdf Restricted to Repository staff only Download (2MB) | Request a copy |
Abstract
Review tentang suatu produk dapat mempengaruhi keputusan pembeli untuk membeli produk tersebut. Semakin ketatnya persaingan penjualan di toko online membawa dampak buruk bagi penulisan review produk. Contoh nyata dialami oleh Amazon.com yang kebanjiran penulis review palsu dan berusaha menuntut para penulis tersebut pada bulan Oktober 2015. Jumlah penulis review palsu tersebut mencapai lebih dari 1000 orang. Menurut Amazon, review produk tersebut “palsu, keliru, tidak otentik, dan kebanyakan rating produk tersebut bintang 5”. Selain mempengaruhi keputusan pembeli, review palsu juga dapat mengganggu pembeli yang mencari informasi dari review jujur dan asli. Oleh karena itu, pada penelitian ini dibangun suatu sistem untuk mendeteksi review spam agar dapat menjadi rekomendasi sistem filter spam otomatis guna mengurangi pengaruh tidak baik pada penjualan toko online maupun pada penulisan review. Ada 3 tipe spam review produk yang biasanya ditemukan, yaitu untruthful spam, brand only spam, dan non review. Tipe pertama sangat sulit dideteksi secara manual. Reviewer bisa saja tidak langung menduplikasi opini, tetapi melakukan menyamaran dengan mengarang kata kata agar terlihat tidak sama. Shingling method dipakai untuk mendeteksi tipe pertama ini. Dari batasan yang ditentukan akan dapat dibedakan review mana yang ber duplikat ataupun near-duplikat sehingga dikategorikan sebagai spam. Untuk tipe 2 dan 3 mendapatkan label awal melalui pelabelan manual. Dari dataset tersebut, 21 review centric feature dan bigram diekstrak dan dilatih menggunakan support vector machine. Untuk pendeteksian tipe 2 dan 3 dilakukan pendekatan self-training semi-supervised learning dengan pertimbangan karena untuk melakukan pelabelan manual memerlukan banyak waktu dan usaha. Hasil akurasi terbaik untuk tipe spam 1 diperoleh dengan praproses tanpa stemming, fitur bigram, dan kernel SVM linear yaitu sebesar 93%. Hasil akurasi terbaik untuk tipe spam 2 dan 3 diperoleh dengan praproses tanpa stemming, penggabungan review centric features dan bigram, oversampling SMOTE borderline1 dan kernel SVM Polynomial yaitu sebesar 86.33%.
===================================================================================================================================
A review of a product can influence the buyer's decision to buy the product. Tight sales competition in online stores has a negative impact on product review writing. A real example experienced by Amazon.com which was flooded with fake review writers and tried to sue the authors in October 2015. The number of fake review writers reached more than 1000 people. According to Amazon, the product review is "fake, wrong, not authentic, and most of the product ratings are 5 star". Fake and spam review also can disturb buyer’s intention to look up information in honest and authentic review. Therefore, this study a system was developed to detect spam reviews so that it could be a recommendation for automated spam filter systems to reduce the level of bad influence on online store sales and product review. There are 3 types of product review spam that are usually found, namely untruthful spam, brand only spam, and non-reviews. The first type is very difficult to detect manually. Reviewers can not directly duplicate opinions, but do disguise by composing words to make them look not the same. The Shingling method is used to detect this first type. From the prescribed limits, it will be able to recognize which reviews are duplicate or near-duplicate so that they are categorized as spam. For types 2 and 3 get the initial label through manual labeling. From the dataset, 21 review centric features and bigram were extracted and trained using the support vector machine. The detection of type 2 and 3 use a semi-supervised learning self-training approach, due to doing manual labeling requires a lot of time and effort. The best accuracy results for type 1 spam were obtained by pre-processing without stemming, bigram feature, and linear SVM kernel which is 93%. The best accuracy results for spam type 2 and 3 were obtained by pre-processing without stemming, merging review centric feature and bigram, SMOTE oversampling borderline1 and polynomial SVM kernel which is equal to 86.33%.
| Item Type: | Thesis (Other) |
|---|---|
| Additional Information: | RSIf 006.312 Zaq d-1 2019 |
| Uncontrolled Keywords: | Opini Spam, Review Spam, Review Centric Features, Metode Shingling, bigram, Oversampling SMOTE, Semi Supervised Learning, SVM |
| Subjects: | T Technology > T Technology (General) > T57.5 Data Processing T Technology > T Technology (General) > T58.5 Information technology. IT--Auditing |
| Divisions: | Faculty of Information and Communication Technology > Informatics > 55201-(S1) Undergraduate Thesis |
| Depositing User: | Ana Alimatus Zaqiyah |
| Date Deposited: | 23 Jul 2026 07:44 |
| Last Modified: | 23 Jul 2026 07:44 |
| URI: | http://repository.its.ac.id/id/eprint/65900 |
Actions (login required)
![]() |
View Item |
