MARATTO

article

Enhancing Arabic Named Entity Recognition with CAMeLBERT, CRF, and Data Augmentation

Abstract

Arabic Named Entity Recognition (NER) remains challenging due to the language’s complex morphology and limited annotated data. This paper proposes an enhanced NER approach that addresses these challenges by combining targeted data augmentation, a language-specific transformer model, and a sequence optimization layer. We expand a small Arabic NER corpus by over 200% using carefully designed augmentation techniques, then fine-tune a specialized Arabic BERT model (CAMeLBERT) with a Conditional Random Field (CRF) decoding layer. Our approach achieved an F1-score of 86.7% on the validation set, substantially outperforming baseline models. The results demonstrate that strategic augmentation coupled with an Arabic-optimized transformer and CRF yields state-of-the-art performance for Arabic NER. This work provides a detailed technical rationale for each design choice and includes a comparative analysis against other deep learning NER architectures to underscore the effectiveness of the proposed method.

Research topics

  • Topic Modeling
  • Natural Language Processing Techniques
  • Authorship Attribution and Profiling

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/cist65886.2025.11224216

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.