Wardana, Bintang Ryan (2026) Deteksi Online Recruitment Fraud (ORF) Menggunakan IndoBERT Dengan Feature Fusion Dan Explainable AI. Other thesis, Institut Teknologi Sepuluh Nopember.
|
Text
5027221022-Undergraduate_Thesis.pdf - Accepted Version Restricted to Repository staff only Download (3MB) | Request a copy |
Abstract
Berdasarkan laporan SEEK, Indonesia menyumbang sekitar 62% dari total kasus Online Recruitment Fraud (ORF) di Asia selama periode Juli 2024 hingga Juni 2025. Tingginya angka tersebut menunjukkan urgensi deteksi ORF secara otomatis, namun pendekatan klasifikasi konvensional umumnya bersifat black-box sehingga sulit menjelaskan alasan di balik suatu keputusan, sementara pendekatan berbasis teks berbahasa Indonesia masih terbatas dieksplorasi. Oleh karena itu, penelitian ini bertujuan mengembangkan dan mengevaluasi model deteksi ORF berbasis IndoBERT yang dikombinasikan dengan feature fusion dan Explainable Artificial Intelligence (XAI) untuk menghasilkan deteksi yang akurat sekaligus transparan. Model yang diusulkan memanfaatkan lima komponen utama teks lowongan kerja, yaitu posisi pekerjaan, profil perusahaan, deskripsi pekerjaan, persyaratan, dan benefit, serta sebelas fitur kontekstual tambahan yang diintegrasikan melalui metode feature fusion. Penelitian ini menggunakan data dari Employment Scam Aegean Dataset (EMSCAD) yang telah diterjemahkan ke dalam bahasa Indonesia. Evaluasi performa dilakukan dengan membandingkan pendekatan feature fusion terhadap pendekatan text-only yang hanya memanfaatkan representasi fitur teks. Performa terbaik dicapai oleh model IndoBERT varian IndoNLU pada skenario text-only yang dilatih dengan proporsi data fraud sebesar 20%, dengan nilai F1-Score 0,883; akurasi 0,950; dan recall 0,936. Sementara itu, pendekatan feature fusion pada varian yang sama menghasilkan F1-Score 0,868; akurasi 0,943; dan recall 0,929. Hasil ini mengindikasikan bahwa representasi teks saja sudah cukup informatif untuk mendeteksi ORF, sedangkan penambahan fitur kontekstual melalui feature fusion belum memberikan peningkatan performa yang signifikan. Selain itu, implementasi XAI menggunakan SHAP dan Integrated Gradients mampu menjelaskan kontribusi setiap komponen teks serta token yang memengaruhi hasil deteksi, sehingga meningkatkan interpretabilitas model.
=================================================================================================================================
According to a SEEK report, Indonesia accounted for approximately 62% of total Online Recruitment Fraud (ORF) cases in Asia during the period of July 2024 to June 2025. This high figure highlights the urgency of automated ORF detection, as conventional classification approaches are generally black-box in nature, making it difficult to explain the reasoning behind their decisions, while approaches based on Indonesian language text remain limited. Therefore, this study aims to develop and evaluate an ORF detection model based on IndoBERT, combined with feature fusion and Explainable Artificial Intelligence (XAI), to achieve detection that is both accurate and transparent. The proposed model utilizes five main components of job vacancy text, namely job title, company profile, job description, requirements, and benefits, along with eleven additional contextual features integrated through the feature fusion method. This study uses data from the Employment Scam Aegean Dataset (EMSCAD), which has been translated into Indonesian. Performance evaluation was conducted by comparing the feature fusion approach with the text-only approach, which relies solely on textual feature representation. The best performance was achieved by the IndoNLU variant of IndoBERT under the text-only scenario trained with a 20% fraud proportion, attaining an F1-Score of 0.883; an accuracy of 0.95; and a recall of 0.936. Meanwhile, the feature fusion approach using the same variant achieved an F1-Score of 0.868; an accuracy of 0.943; and a recall of 0.929. These results indicate that text representation alone is already sufficiently informative for detecting ORF, while the addition of contextual features through feature fusion has not yet provided a significant improvement in performance. In addition, the implementation of XAI using SHAP and Integrated Gradients was able to explain the contribution of each text component and token influencing the detection results, thereby improving the interpretability of the model.
| Item Type: | Thesis (Other) |
|---|---|
| Uncontrolled Keywords: | Online Recruitment Fraud, IndoBERT, Feature Fusion, Deep Learning, Explainable AI |
| Subjects: | T Technology > T Technology (General) > T57.84 Heuristic algorithms. |
| Divisions: | Faculty of Intelligent Electrical and Informatics Technology (ELECTICS) > Information Technology > 59201-(S1) Undergraduate Thesis |
| Depositing User: | Bintang Ryan Wardana |
| Date Deposited: | 30 Jul 2026 07:30 |
| Last Modified: | 30 Jul 2026 07:30 |
| URI: | http://repository.its.ac.id/id/eprint/139927 |
Actions (login required)
![]() |
View Item |
