MARATTO

article · Journal of King Saud University - Computer and Information Sciences

AlgVec: A word embedding model for the algerian dialect in arabic and arabizi

Abstract

This paper introduces AlgVec, a suite of word embedding models trained specifically for the Algerian dialect, a linguistically rich but under-resourced variety of Arabic. The embeddings are derived from a large corpus of user-generated content on the social media platform X (formerly Twitter), covering both Arabic script and Arabizi. The Algerian dialect presents unique challenges for Natural Language Processing (NLP), including informal grammatical structures, non-standardized spelling, and frequent code-switching. To address these issues, we compile and preprocess a corpus of more than 32 million tokens and train multiple embedding models using Word2Vec (Skip-gram and Continuous Bag of Words (CBOW)) as well as FastText architectures. We evaluate AlgVec through intrinsic tasks such as word similarity, nearest-neighbor retrieval, and a linguistically grounded DiaLex-style benchmark adapted for Arabic dialects, as well as through a downstream sentiment analysis task. Furthermore, we explore a combined embedding approach that integrates AlgVec with the transformer-based MARBERT model for sentiment analysis, using Support Vector Machine (SVM), Convolutional Neural Network (CNN), and Long Short-Term Memory (LSTM) classifiers. Our experiments show that AlgVec and the combined model outperform widely used Modern Standard Arabic (MSA) embeddings on dialect-specific benchmarks. The complete set of embeddings is publicly released to support future research in Arabic dialect NLP.

Research topics

  • Authorship Attribution and Profiling
  • Linguistic Variation and Morphology
  • Natural Language Processing Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1007/s44443-025-00407-6

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.