article · Revue d intelligence artificielle
Distributed representations of words in a vector space help learning algorithms to model semantic notions of word similarity and distances in sentences.Most of the existing researches have been done on the Latin, Arabic, and other language, while the Amazigh language is ignored.In this paper, we try to build a first model word embeddings for Amazigh language and describe the steps needed to build it.Therefore, we implement a Word2Vec that a combination of two techniques -CBOW (Continuous bag of words) and Skip-gram to transform words written in Tifinagh to vector form.To obtain the highest performance, we evaluate two parameters of Word2Vec include Word2Vec model architecture and vector dimension.This evaluation process was implemented towards our proposed corpus collected on Amazigh websites for different domains.The result shows that the highest accuracy values are obtained under the combination of CBOW model and 300 dimensional vector.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.18280/ria.370324
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.