MARATTO

article

Disambiguating Setswana conjunctions with BERT-based models

Abstract

Word Sense Disambiguation (WSD) presents significant challenges in natural language processing, particularly for under-resourced languages such as Setswana. This study evaluates six advanced language models on their ability to disambiguate multiple senses of common Setswana conjunctions, employing accuracy, F1-score, and Quadratic Weighted Kappa (QWK) as evaluation metrics. The findings reveal that LaBSE achieved the highest overall scores in simpler contexts with fewer senses, peaking at an accuracy of 83.00% and a QWK of 66.00% for the conjunction ”mme.” In contrast, PuoBERTa, while optimized for Setswana, excelled in more complex scenarios involving conjunctions with multiple senses, underscoring the importance of model choice based on the linguistic complexity of the task. <br/><br/> These results emphasize the critical role of tailored language models in enhancing WSD tasks for under-resourced languages. They demonstrate that specific adjustments to model training and architecture can significantly improve performance, thereby increasing the precision and applicability of NLP technologies in diverse linguistic settings. This research not only augments computational resources for Setswana but also provides a blueprint for applying similar methodologies to other less-represented languages, advancing global communication technologies.

Research topics

  • Semantic Web and Ontologies
  • Model-Driven Software Engineering Techniques
  • Service-Oriented Architecture and Web Services

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1117/12.3050015

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.