article
The scarcity of labeled datasets poses significant challenges in Natural Language Processing (NLP), particularly for sentiment and partiality analysis of Arabic tweets regarding the Russo-Ukrainian War. Semi-supervised learning (SSL) addresses this issue by leveraging both labeled and unlabeled data. This paper investigates SSL approaches in Arabic NLP by analyzing sentiment and partiality in collected tweets. We trained four models, each with a different approach: (1) a supervised model, (2) a model initialized with pretraining, (3) a model trained with consistency regularization, and (4) a model combining pretraining and consistency regularization. We limited our dataset to 200,000 labeled to compare the effectiveness of the chosen SSL methods with previous work done on a 500,000 tweet dataset. Then we built further over the best model obtained from the previous ones by training it model on 1.3 million labeled tweet, to acheive a state-of-the-art performance 98.9% test accuracy. Our findings demonstrate that SSL can reduce the dependence on labeled data while still achieving competitive performance. This study underscores SSL's potential to enhance Arabic NLP applications, providing more efficient and scalable solutions.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/niles63360.2024.10753204
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.