article
Data representation plays a crucial role in machine learning tasks, especially with the rise of transformer models. While many studies focus on identifying the best-performing model, there is a need to understand how altering vector representations affects these models. This research paper delves into modifying AraBERT vectors to categorize data into 26 distinct classes. Our primary data sources include a multi-dialect dataset curated for dialect identification tasks. This dataset comprises comments collected and annotated as part of a shared task, resulting in a sizable dataset of over 54,000 annotated comments. Our system utilizes multiple vector inputs to transformer models, a class of deep learning models renowned for their efficacy in natural language processing tasks, leveraging their combined capabilities. The obtained results show an improvement of 17 points F-measure.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/iscv60512.2024.10620152
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.