MARATTO

dataset · Mendeley Data

RIYE Audio Dataset: A Multidialectal Speech Corpus for Low-Resource Language Processing

Abstract

This dataset consists of a curated collection of high-fidelity, field-recorded audio samples developed under the Digiculture RIYE Project's ethnographic survey framework. Created to bridge the digital divide for under-resourced languages, this corpus captures diverse regional speech patterns, unique tonal variations, and distinct phonetic markers essential for localised speech research. The dataset is designed to support a wide array of machine learning, signal processing, and computational linguistics tasks. Because the audio is formatted for clean feature engineering, it serves as an ideal baseline for researchers developing neural networks, automatic speech recognition (ASR) engines, and lightweight audio classification systems optimised for edge deployment.

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.17632/kt996wpns5

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.