article
The performance of learning models mostly depends on the availability and quality of training data. To address the issue of dataset adequacy, researchers have widely investigated Data Augmentation (DA) as a promising solution. This paper examines the application of data augmentation techniques for the Moroccan Darija dialect, an under-resourced language with limited linguistic resources. We investigate using four Easy Data Augmentation (EDA) methods—Random Swap, Random Deletion, Random Insertion, and Synonym Replacement—to increase the diversity of training data, ultimately improving the performance of machine learning models. These techniques were applied to a corpus of Moroccan Darija sentences to improve text classification tasks for emotion detection (Fear, Anger, and Joy). The results demonstrate that these augmentation techniques can significantly enhance data variability, leading to more robust models for the Darija dialect. By creating synthetic examples, this approach tackles the issue of data scarcity in Moroccan Darija, ultimately improving performance in Natural Language Processing (NLP) tasks for low-resource languages.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset64571.2025.11007959
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.