MARATTO

article

Parallel Corpora Preparation for bi-Directional Amharic-Kistanigna Machine Translation

Abstract

This research focused on creating a parallel corpus for the Amharic and Kistanigna languages and conducting Machine Translation (MT) tests on it. The goal was to increase Kistanigna language content online, address its endangered status, and facilitate information sharing between the two languages. To achieve this, a parallel corpus was developed, and machine translation models were tested, including LSTM, BiLSTM, LSTM with attention, CNN with attention, and Transformer models. Experiments were conducted with both word and morpheme-based translation units, using the morfessor tool for morpheme segmentation. The best results were achieved with morpheme-based bidirectional machine translation using a Transformer, with BLEU scores of 21.31 for Amharic-Kistanigna and 22.40 for Kistanigna-Amharic translations. The resulting corpus contains 9, 225 parallel sentences and will be made available to the community. This is the first parallel corpus for these languages, and more such corpora are needed for further research.

Research topics

  • Natural Language Processing Techniques
  • Language, Linguistics, Cultural Analysis
  • Handwritten Text Recognition Techniques

Sustainable Development Goals

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1109/ict4da62874.2024.10777170

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.