article
this study examines the difficulties of classifying text in Arabic using advanced machine learning (ML) algorithms with dimension reduction methods like Principal Component Analysis (PCA) and Uniform Manifold Approximation and Projection (UMAP). Keeping in view the unique complexities of text in news articles written in Arabic, we employed diverse ML algorithms like Support Vector Machines (SVM), Logistic Regression (LR), Multinomial Naïve Bayes (MNB), and Random Forest. In comparative research, we examine the effect of PCA and UMAP on model performance with regard to accuracy and processing time. The findings indicate that PCA increases accuracy in all models with a maximum accuracy of 87.23% using SVM with PCA. Along with this, PCA reduces processing time significantly compared to processing raw text and is thus a good candidate to consider in text classification in large datasets. This research not only emphasizes the importance of dimension reduction in text classification in Arabic but also offers insights to enhance ML workflows in other languages with complex structures.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iraset64571.2025.11008216
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.