article
Cyberbullying is an increasing threat on social media, with serious consequences for mental health, particularly among young people. Despite growing global efforts to address this problem, developing accurate detection systems remains challenging for low-resource languages in both text and audio modalities, such as Arabic and its dialects. This paper presents a binary-labeled dataset for cyberbullying detection in Moroccan Darija. The dataset merges over 4,000 newly collected YouTube comments with the OMCD (Offensive Moroccan Comments Dataset) corpus of 8,024 comments, originally labeled for offensive language. To better reflect the nature of cyberbullying, the entire corpus was reannotated from scratch, clearly distinguishing between bullying and non-bullying content, including subtle forms like sarcasm, shaming, and indirect aggression. Several Arabic transformer models were fine-tuned and evaluated using standard classification metrics. The results show that dialectspecific models, particularly DarijaBERT, achieve the best performance, underlining the importance of context-aware annotation and in-domain pre-training for cyberbullying detection in lowresource settings.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/wincom65874.2025.11313421
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.