article · Journal of King Saud University - Computer and Information Sciences
Arabic Part of Speech tagging faces significant challenges because contemporary texts largely omit diacritics and Arabic derivatives are difficult to differentiate. Consequently, the same written form can represent multiple words, requiring advanced processing resources to assign accurate grammatical tags. To address this, a structured framework provides detailed hierarchical levels of Arabic tagset categories and their relationships. Grounded in a comparative study and established Arabic grammar references, this hierarchical design enables straightforward expansion while delivering more precise and accurate tagging outcomes. Subject-matter experts validated the categories, after which the proposed tagset was implemented within a Part of Speech tagger. Subsequent experimental testing demonstrated the performance of the system. The developed tagset marks progress towards establishing a standardised, rich, and comprehensive Part of Speech tagging system for Arabic language processing.
Arabic natural language processing tools struggle with ambiguity caused by the absence of diacritics in everyday written text. Developing a standardised, hierarchical tagset validated by linguistic experts provides a structured way to handle complex word derivations. This improves the accuracy of grammatical analysis, helping computers better interpret Arabic text for downstream language processing tasks.
The work could enhance natural language processing software, search engines, and automated text analysis tools handling modern Arabic text. Software developers and language technology companies could use the structured tagset to improve the accuracy of their parsing systems. Because the tagset has been implemented in a tagger and evaluated through experimental testing, it appears to be at an applied and tested stage of development.
AI-generated from the published abstract. Always read the original work before citing.
Part of Speech (PoS) tagging is still not very well investigated with respect to the Arabic language. Determining the PoS tags of a word in a particular context is difficult, primarily because there is no use of diacritics in most of contemporary texts. Consequently, the same word may be spelled in different ways. Further, detecting the difference between Arabic derivatives represents a very challenging issue for the majority of PoS taggers. Hence, the task of tagging the correct PoS tags requires advanced processing and the use of considerable resources. This study aims to design detailed hierarchical levels of the Arabic tagset categories and their relationships. These hierarchical levels allow easier expansion when required and produce more accurate and precise results. They are based on a comparative study and important references in Arabic grammar; they are also validated by experts in this field. In addition, the proposed tagset is implemented in a PoS tagger and tested via various experiments. We believe that our study makes a significant contribution to the literature because this work is an advancement in the direction of achieving a standard, rich, and comprehensive tagset for Arabic.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1016/j.jksuci.2017.01.006
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.