MARATTO

article

A BERT Deep Learning Model for Arabic Spam Detection

Abstract

Spam messages pose a significant cybersecurity threat, leading to phishing attacks, fraud, and privacy breaches. Traditional spam detection methods, such as rule-based filtering and statistical models, often fail to capture the evolving and complex nature of spam messages. In this paper, we propose an Arabic spam detection model leveraging BERT (Bidirectional Encoder Representations from Transformers), a deep learning-based NLP model. Our approach enhances classification accuracy by utilizing contextual text representations specific to the Arabic language. We preprocess Arabic text using AraBERT tokenization and fine-tune the BERT-based model on a balanced dataset of Arabic spam and ham messages. Experimental results demonstrate that our model achieves high accuracy (98%), outperforming traditional machine learning and deep learning approaches. This research highlights the potential of transformer-based models in Arabic spam filtering, paving the way for more efficient and robust detection systems.

Research topics

  • Spam and Phishing Detection
  • Text and Document Classification Technologies
  • Imbalanced Data Classification Techniques

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/codit66093.2025.11321615

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.