dataset · Zenodo (CERN European Organization for Nuclear Research)
This release presents a human-validated Phase-2 expansion of the Fusha–Darija evaluation project, offering a comprehensive dataset for linguistic research. It includes publishable experiment materials and anonymised human annotations across several linguistic dimensions. These dimensions cover lexical semantics, etymological evidence, Modern Standard Arabic–Moroccan Darija semantic pairs, comprehension tasks, cross-variety Arabic counterparts, framing, register continua, script and code-switching variation, and digital-service requests. The dataset features a public reference layer that preserves detailed annotation statuses, rather than simplifying uncertainty. Comprising 28 validated data, annotation, and reference tables with 11,284 rows, this resource supplements an earlier Harvard Dataverse package.
This dataset provides valuable, human-validated linguistic data for understanding the relationship between Modern Standard Arabic (Fusha) and Moroccan Darija. It supports research into how these Arabic varieties are used and understood, which is crucial for developing language technologies, educational resources, and improving communication across different Arabic-speaking contexts.
This linguistic dataset serves as foundational research infrastructure. It could be utilised by researchers and developers working on natural language processing tools, machine translation, or educational software tailored for Arabic varieties. The focus on semantic pairs, code-switching, and digital-service requests suggests potential for improving AI models that interact with users in these languages. This appears to be early-stage research infrastructure, providing data for future applied developments.
AI-generated from the published abstract. Always read the original work before citing.
Human-validated Phase-2 expansion of the Fusha–Darija evaluation project. The release contains publishable experiment materials and anonymized human annotations for E01, E02, E03, E04, E05, E08, and E09, covering lexical semantics and etymological evidence, Modern Standard Arabic–Moroccan Darija semantic pairs and comprehension tasks, cross-variety Arabic counterparts, framing, register continua, script and code-switching variation, and digital-service requests. The public reference layer preserves primary-gold, multi-reference, ambiguous, stress-condition, and excluded-from-primary statuses rather than collapsing uncertainty. The release contains 28 validated data/annotation/reference tables with 11,284 rows. Raw reviewer workbooks, reviewer identities, private communications, model outputs, and publication results are excluded. This dataset supplements and does not replace the Harvard Dataverse V1 package: https://doi.org/10.7910/DVN/W0C5DV. The archive is mixed-license: project-original materials and derived anonymized annotations are CC BY 4.0; selected third-party components retain CC BY-NC 4.0 or CC BY-SA 4.0 as documented at component and row level within the archive.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.5281/zenodo.21580964
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.