Comparison of Support Vector Machine (SVM) and Random Forest Algorithm for Detection of Negative Content on Websites
DOI:
https://doi.org/10.26555/jiteki.v9i1.25861Keywords:
Negative Content, Natural Language Processing, Machine Learning, Support Vector Machine, Random ForestAbstract
The amount of negative content circulating on the internet can damage people's morale so that social conflicts arise in society that threaten national sovereignty. Detecting negative content can help identify and prevent harmful events before they occur. This can lead to a safer and more positive online environment. Comparison of Support Vector Machine (SVM) and Random Forest (RF) Algorithm for Detection of Negative Content on Websites. The research contributions are 1) detect negative content on the internet with random forest and SVM, 2) comparing SVM and RF algorithms for detecting negative content on websites, 3) detection of negative content based on text focusing on the categories of fraud, gambling, pornography and Whitelist. The stages of this research are preparing a text content dataset on a website that has been labeled, preprocessing (duplicated data, text cleansing, case folding, stopward, tokenize, label encoding, data splitting, and determine the TF-IDF), finally performing the classification process with SVM and Random Forest. The dataset used in this study is a structured dataset in the form of text obtained from emails that have been registered on the TrustPositive website as negative content. Negative content includes fraud, pornography and gambling. The results show the accuracy of the SVM is 97%, Precision 90% and Recall 91%, while for Accuracy in Random Forest is 92%, Precision 71%, and Recall 86%. The value obtained is the result of testing using 526 website URLs. The test results show that the Support Vector Machine is better than the Random Forest in this study.Downloads
Published
2023-03-20
Issue
Section
Articles
License
Authors who publish with JITEKI agree to the following terms:
- Authors retain copyright and grant the journal the right of first publication with the work simultaneously licensed under a Creative Commons Attribution License (CC BY-SA 4.0) that allows others to share the work with an acknowledgment of the work's authorship and initial publication in this journal.
- Authors are able to enter into separate, additional contractual arrangements for the non-exclusive distribution of the journal's published version of the work (e.g., post it to an institutional repository or publish it in a book), with an acknowledgment of its initial publication in this journal.
- Authors are permitted and encouraged to post their work online (e.g., in institutional repositories or on their website) prior to and during the submission process, as it can lead to productive exchanges, as well as earlier and greater citation of published work.
This work is licensed under a Creative Commons Attribution 4.0 International License