MARATTO

article

Alkhalil Corpus: An Open-Source Thematic and Lemmatized Corpus for Modern Standard Arabic

2026Open accessMohamed I University

Abstract

The availability of large annotated corpora remains a major challenge for the development of natural language processing systems for underresourced languages such as Arabic.In this paper, we present two annotated corpora dedicated to Modern Standard Arabic.These corpora are open-source and freely available on the Hugging Face platform.The first corpus, annotated by theme and designed to provide a balanced representation of contemporary Arabic usage, comprises approximately 76 million words collected from diverse sources covering multiple domains and geographical regions.The second corpus, containing approximately one million words, is a sub-corpus extracted from the first.It was annotated with lemma tags using a semi-automatic approach that combines automatic annotation with the Alkhalil lemmatizer and MADAMIRA, followed by manual validation.

Research topics

  • Language, Linguistics, Cultural Analysis
  • Natural Language Processing Techniques
  • Medieval and Classical Philosophy

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.18653/v1/2026.abjadnlp-1.27

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.