MARATTO

dataset · Zenodo (CERN European Organization for Nuclear Research)

Runyankore-Rukiga Simulated Speech Corpus for Cross-Lingual Computational Speech Research (Version 1.0)

Abstract

The Runyankore-Rukiga Simulated Speech Corpus (RR-SC) is a structured speech dataset developed to support computational speech analysis and cross-lingual methodological research in a low-resource African language. The corpus was constructed through translation and scripted enactment of standardized Cookie Theft picture-description transcripts derived from the DementiaBank Pitt Corpus. English transcripts were translated into Runyankore-Rukiga by bilingual native speakers using a structured annotation framework that preserved pauses, hesitations, and paralinguistic markers. Independent native-speaker review was performed to ensure semantic fidelity and annotation consistency. The dataset contains 481 paired audio recordings and corresponding transcripts generated by healthy adult native Runyankore-Rukiga speakers under controlled recording conditions. Of these recordings, 216 were produced under a fluent speech condition and 265 under a reduced-fluency simulated speech condition. The terms Control for Fluent speech condition and Dementia for reduced-fluency simulated speech condition were maintained in the labelling of the audios and transcripts in order to preserve the audio originality. Audio recordings were collected using a Xiaomi Redmi Note 13 Pro smartphone at 44.1 kHz sampling rate and 16-bit resolution and stored in WAV format. Transcripts were stored in CHAT (.cha) format and annotated using a reproducible symbol-based annotation scheme that explicitly encodes pauses, noisy pauses, laughter events, confused speech, and accelerated speech segments. The corpus was designed as a methodological and computational resource for: • Cross-lingual speech processing research • Low-resource language speech research • Feature extraction benchmarking • Speech representation learning • Computational reproducibility studies • Educational and training applications Important Notice: This dataset does not contain recordings from individuals clinically diagnosed with dementia or cognitive impairment. All recordings were produced by healthy adult volunteers performing scripted enactments of translated narrative transcripts. The dataset is intended solely for methodological experimentation, computational speech research, benchmarking, and educational purposes. It must not be used for clinical diagnosis, dementia screening, biomarker discovery, or medical decision-making. Ethical approval was obtained from the Mbarara University of Science and Technology Research Ethics Committee (MUST-2024-1723) and the Uganda National Council for Science and Technology (HS5749ES). Participants provided informed consent for research use and public scientific dissemination of their recordings.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.5281/zenodo.20478456

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.