article
The use of technology products and smart devices in Morocco has increased in recent years, leading to the production of a large amount of data. This data can be used to provide users with personalized experiences and recommendations, to help companies understand consumer behavior and market trends, and to support research and innovation in various fields. Processing Moroccan text data is not an easy task due to many challenges facing this dialect, including characteristics of the language, the lack of data resources, different writing scripts and many more. That being said, the aim of this paper is to present a pipeline of building a large annotated biscript lexicon for Moroccan Dialect with both Arabizi and Arabic. A description of the collection and processing steps from cleaning data to tokenization to filtering and refining were presented. The paper also outlines the annotation and the validation process of this proposed approach.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/wincom59760.2023.10322961
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.