MARATTO

article

Enhancing Automatic Speech Recognition for Moroccan Darija Using Transformer Models

Abstract

Automatic Speech Recognition (ASR) for the Moroccan Darija a dialect characterized by high linguistic variability and a lack of standardization is a challenging task. We fine-tuned two state-of-the-art transformer-based models, Wav2Vec2-XLSR-53 and Whisper, and created a 5-hour high-quality Moroccan Darija corpus extracted from public YouTube channels. The results show that, in noisy conditions, Whisper is more accurate; however, Wav2Vec2 outperforms Whisper in clean conditions, reaching a Word Error Rate (WER) of 10.64% and a Character Error Rate (CER) of 2.22%. This work can help determine how to set up an effective ASR system for Darija and other low-resource, heavily code-switching dialects.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/sita67914.2025.11273609

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.