review · Technologies
This systematic review examined advances in neural machine translation (NMT) and large language models (LLMs) for low-resource languages between 2017 and 2025. The study, following PRISMA guidelines, analysed 63 articles to address the widening gap in AI-driven language technologies for marginalised languages. It identified five key methodological approaches, including data augmentation and transfer learning. Findings indicate that model performance varies with resource availability: transformer-based NMT performs well with moderate data, while LLMs show promise in extremely low-resource settings. Hybrid NMT–LLM approaches were found to be particularly effective. The review also highlighted critical challenges such as a lack of standardised benchmarks, over-reliance on inadequate evaluation metrics, and ethical concerns.
This research is important because it addresses the disparity in AI language technologies, which currently favour high-resource languages. By identifying effective methods and critical challenges for low-resource translation, it helps advance inclusive and equitable AI, ensuring that more communities can benefit from technological progress and participate in the digital economy.
The findings could inform the development of more effective machine translation tools for businesses and organisations operating in regions with diverse, low-resource languages. Potential users include technology companies, international aid organisations, and government bodies needing to communicate across language barriers. This research is foundational, identifying effective approaches and challenges, suggesting it is at an early-stage research level, guiding future applied development.
AI-generated from the published abstract. Always read the original work before citing.
The rapid evolution of AI-driven language technologies has inadvertently widened the gap between high-resource and marginalised languages. Despite significant progress in AI-driven translation for high-resource languages, low-resource languages remain underrepresented due to limited data, a lack of benchmarks, and evaluation challenges. This study presents a comprehensive systematic review of machine translation for low-resource languages, focusing on advances in neural machine translation (NMT) and large language models (LLMs) between 2017 and 2025. Following Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines, 63 studies were selected from the 1696 articles in the Scopus, Web of Science, and Google Scholar databases. The review identifies five dominant methodological approaches: data augmentation, back-translation, transfer learning, pre-training, and parameter-efficient fine-tuning. The findings reveal that model performance is highly dependent on resource availability: transformer-based NMT excels in moderate data settings, while LLMs demonstrate promising zero-shot and few-shot capabilities in extremely low-resource scenarios. Hybrid NMT–LLM approaches emerge as a particularly effective paradigm. The study also highlights critical challenges, including the absence of standardised benchmarks, over-reliance on inadequate evaluation metrics such as Bilingual Evaluation Understudy (BLEU), limited human evaluation, and significant geographic and linguistic underrepresentation. Additionally, ethical concerns related to bias, cultural representation, and community engagement are increasingly relevant. The findings contribute to advancing inclusive and equitable AI-driven language technologies.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.3390/technologies14080518
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.