MARATTO

article

A Proposed Approach for Extracting Semantic and Lexical Relations for Low-Resource Languages: A Case Study of Darija

Abstract

Extracting semantic relations between words is crucial for the development and enrichment of lexical resources, especially for under-resourced languages like Moroccan Darija. This paper presents an automated methodology for identifying synonyms, antonyms, hypernyms, and hyponyms by leveraging bilingual Darija–English resources, Princeton WordNet (PWN), the Suggested Upper Merged Ontology (SUMO), and the NLTK toolkit. Experimental evaluation was conducted on a dataset of 361 Darija nouns, selected as a preliminary testbed to validate the methodology before scaling it to the full lexicon. The results show that 83.10% were successfully aligned with PWN synsets, resulting in the extraction of 14,201 semantic relations, of which 5,475 (38.55%) were validated through back-translation. These findings confirm the potential of transferring semantic knowledge from English into Darija, despite cultural and lexical mismatches. The proposed pipeline substantially enriches Darija’s lexical coverage and offers a scalable and replicable approach for developing semantic resources in other low-resource dialects.

Research topics

  • Natural Language Processing Techniques
  • Language and cultural evolution
  • linguistics and terminology studies

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/cist65886.2025.11224229

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.