MARATTO

article · East African Journal of Information Technology

KenLumachiQuAD – A Question Answering Dataset for Kenyan Luhya Lumarachi Language for Machine Learning

2026Open accessUniversity of Nairobi

Abstract

Question-answering (QA) datasets play a crucial role in testing and training machine learning models, from which we can develop practical end-user applications, such as internet search, dialogue systems, and chatbots. There are various methods for creating QA datasets, including the use of transfer learning, employing translations, creation from synthetic data, and the use of human annotators, depending on the availability of reference datasets. QA datasets for low-resource languages are few and tend to be of small sizes due to the costs associated with human annotation, which is the easiest available method for such low-resource languages in the absence of other methods that need some existing reference data. We make our contribution in the provision of QA resources for low-resource languages of Kenya by developing a human-annotated QA dataset, called KenLumachiQuAD. This is a QA dataset for the Kenyan low-resource language of Luhya, specifically the Lumarachi dialect. KenLumachiQuAD is a dataset of 1,000 QA pairs that is human-annotated from available public domain texts for the Luhya Lumarachi language. We evaluated the dataset on its applicability to the machine learning task of QA using a semantic network modelling method, based on a small sample of the dataset, and achieved a result of 76% exact match. The research, therefore, provides a QA dataset for machine learning and contributes to the resourcing of low-resource languages. Researchers can still add more QA to this dataset from texts that are yet to be annotated to further augment this dataset. Finally, other researchers can apply our experience in developing other QA datasets for low-resource languages

Research topics

  • Natural Language Processing Techniques
  • Topic Modeling
  • Language and cultural evolution

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.37284/eajit.9.1.4519

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.