Comparison of Support Vector Machine (SVM) and Random Forest Algorithm for Detection of Negative Content on Websites

Hermawan Syahputra, Aldiva Wibowo

Abstract


The amount of negative content circulating on the internet can damage people's morale so that social conflicts arise in society that threaten national sovereignty. Detecting negative content can help identify and prevent harmful events before they occur. This can lead to a safer and more positive online environment. Comparison of Support Vector Machine (SVM) and Random Forest (RF) Algorithm for Detection of Negative Content on Websites. The research contributions are 1) detect negative content on the internet with random forest and SVM, 2) comparing SVM and RF algorithms for detecting negative content on websites, 3) detection of negative content based on text focusing on the categories of fraud, gambling, pornography and Whitelist. The stages of this research are preparing a text content dataset on a website that has been labeled, preprocessing (duplicated data, text cleansing, case folding, stopward, tokenize, label encoding, data splitting, and determine the TF-IDF), finally performing the classification process with SVM and Random Forest. The dataset used in this study is a structured dataset in the form of text obtained from emails that have been registered on the TrustPositive website as negative content.  Negative content includes fraud, pornography and gambling. The results show the accuracy of the SVM is 97%, Precision 90% and Recall 91%, while for Accuracy in Random Forest is 92%, Precision 71%, and Recall 86%. The value obtained is the result of testing using 526 website URLs. The test results show that the Support Vector Machine is better than the Random Forest in this study.

Keywords


Negative Content; Natural Language Processing; Machine Learning; Support Vector Machine; Random Forest

Full Text:

PDF


DOI: http://dx.doi.org/10.26555/jiteki.v9i1.25861

Refbacks

  • There are currently no refbacks.


Copyright (c) 2023 Hermawan Syahputra

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.


 
About the JournalJournal PoliciesAuthor Information
 


Jurnal Ilmiah Teknik Elektro Komputer dan Informatika
ISSN 2338-3070 (print) | 2338-3062 (online)
Organized by Electrical Engineering Department - Universitas Ahmad Dahlan
Published by Universitas Ahmad Dahlan
Website: http://journal.uad.ac.id/index.php/jiteki
Email 1: jiteki@ee.uad.ac.id
Email 2: alfianmaarif@ee.uad.ac.id
Office Address: Kantor Program Studi Teknik Elektro, Lantai 6 Sayap Barat, Kampus 4 UAD, Jl. Ringroad Selatan, Tamanan, Kec. Banguntapan, Bantul, Daerah Istimewa Yogyakarta 55191, Indonesia