article · Language Resources and Evaluation
Online abuse, hate speech, and cyberbullying across social media platforms create a pressing need for automated moderation of harmful content. A balanced Arabic dataset was compiled to assist in detecting abusive material, supplementing two existing publicly available datasets. Three individual machine learning classifiers and three ensemble models were evaluated to determine their ability to detect offensive language and cyberbullying in Arabic text. Across all three evaluated datasets, ensemble machine learning methods outperformed individual classifiers. A voting ensemble emerged as the most accurate model, registering accuracy scores of 71.1 percent, 76.7 percent, and 98.5 percent, compared to top single classifier scores of 65.1 percent, 76.2 percent, and 98 percent. The performance of this voting method was subsequently improved through hyperparameter tuning applied specifically to the Arabic cyberbullying dataset.
The proliferation of harmful content such as hate speech, harassment, and racism across social media platforms presents serious challenges for online safety. Developing automated tools tailored to Arabic text helps regulate offensive interactions and limit abusive material on social networks, offering more dependable automated monitoring than single algorithm approaches.
This work could enable automated moderation tools for social media platforms, communication services, and online communities monitoring Arabic content. Intended users include digital platforms seeking to regulate abusive behaviour, hate speech, and cyberbullying. As an applied and tested algorithmic study evaluating classifier configurations across specific datasets, the models remain at an experimental stage and require integration into operational moderation pipelines before commercial deployment.
AI-generated from the published abstract. Always read the original work before citing.
Abstract Since cyberbullying impacts both individual victims and entire society, research on abusive language and its detection has attracted attention in recent years. Because social media sites like Facebook, Instagram, Twitter, and others are so widely accessible, hate speech, bullying, sexism, racism, aggressive material, harassment, poisonous comments, and other types of abuse have all substantially increased. Due to the critical requirement to detect, regulate, and limit the spread of harmful content on social networking sites, we conducted this study to automate the detection of offensive language or cyberbullying. We created a new Arabic balanced data set to be used in the offensive detection process because having a balanced data set for a model would result in improved accuracy models. Recently, the performance of single classifiers has been improved using ensemble machine learning. The purpose of this study is to examine the effectiveness of several single and ensemble machine learning algorithms in identifying Arabic text that contains foul language and cyberbullying. Applying them to three Arabic datasets, we have selected three machine learning classifiers and three ensemble models for this aim. Two of them are offensive datasets that are readily accessible in the public, while the third one was created. The results showed that the single learner machine learning strategy is inferior to the ensemble machine learning methodology. Voting performs is the best performing trained ensemble machine learning classifier, outperforming the best single learner classifier (65.1%, 76.2%, and 98%) for the same datasets with accuracy scores of (71.1%, 76.7%, and 98.5%) for each of the three datasets used. Finally, we improve the voting technique’s performance through hyperparameter tuning on the Arabic cyberbullying data set.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1007/s10579-023-09683-y
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.