data paper · Data in Brief
This article presents AryWiki-Instruct, a high-fidelity instruction tuning dataset for Moroccan Arabic (Darija), comprising 46,590 Question and Answer (QA) pairs. The dataset was derived from a snapshot of the Moroccan Arabic Wikipedia ( arywiki ) and generated using the Gemini-2.5-Flash model via a Context Aware batch processing architecture. The data creation process involved parsing raw XML Wikipedia dumps, filtering for script consistency and token density, and applying a Context Injection generation strategy to prevent coreference ambiguity. To ensure high information density, the raw generated output (82,500 pairs) was subjected to a rigorous automated quality assurance pipeline. This pipeline utilized a hierarchical trigger confirmation algorithm to remove repetitive administrative census noise and employed composite embedding based clustering ( multilingual-e5-large ) for semantic deduplication. The final dataset is formatted as a JSONL file, providing paired instructions and responses alongside their taxonomic categories. This dataset provides a native first culturally grounded resource designed to facilitate the supervised fine tuning of Large Language Models (LLMs) in Maghrebi dialects, circumventing the syntactic limitations of translated English instruction sets.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.dib.2026.113140
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.