article
Recent global interest in Hadith literature has encouraged efforts to enhance the efficiency of Hadith classification in Arabic. This study investigates the impact of different text representation models on Hadith classification. We introduce a novel model combining two text representation techniques: Term Frequency-Inverse Document Frequency (TF-IDF) and word embedding followed by dimension reduction using Principal Component Analysis (PCA). This approach allows us to emphasize word importance and capture the semantic meaning of each word within the dataset. Our dataset includes 834 hadiths from Sahih Al-Bukhari, representing five distinct categories. The model has been evaluated using various machine-learning algorithms, and the results show that the Stochastic Gradient Descent (SGD) algorithm achieved the highest F1-score ($86.99 \%$) followed by the Multi-Layer Perceptron (MLP) algorithm with (81.53%). These findings indicate the potential of our approach to significantly improve the categorization of Hadith in the Arabic language, which can have a profound impact on research in Islamic studies.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/jac-ecc64419.2024.11061234
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.