article
Automatic Speech Recognition (ASR) for the Moroccan Darija a dialect characterized by high linguistic variability and a lack of standardization is a challenging task. We fine-tuned two state-of-the-art transformer-based models, Wav2Vec2-XLSR-53 and Whisper, and created a 5-hour high-quality Moroccan Darija corpus extracted from public YouTube channels. The results show that, in noisy conditions, Whisper is more accurate; however, Wav2Vec2 outperforms Whisper in clean conditions, reaching a Word Error Rate (WER) of 10.64% and a Character Error Rate (CER) of 2.22%. This work can help determine how to set up an effective ASR system for Darija and other low-resource, heavily code-switching dialects.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/sita67914.2025.11273609
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.