article · ACM Transactions on Asian and Low-Resource Language Information Processing
Many sub-Saharan African languages are categorized as tone languages, and for the most part, they are classified as low-resource languages due to the limited resources and tools available to process these languages. Identifying the tone associated with a syllable is therefore a key challenge for speech recognition in these languages. We propose models that automate the recognition of tones in continuous speech that can easily be incorporated into a speech recognition pipeline for these languages. We have investigated different neural architectures as well as several feature extraction algorithms in speech (FBs (Filter Banks), LEAF (Learnable Frontend), CS (Cestrogram), MFCC (Mel-Frequency Cepstral Coefficients)). In the context of low-resource languages, we also evaluated W2V (Wav2vec 2.0) models for this task. In this work, we use a public speech recognition dataset on Yoruba. As for the results, using the combination of features obtained from CS and FBs, we obtain a minimum TER (Tone Error Rate) of 19.54%, whereas for the evaluations of the models using W2V, we have a TER of 17.72%, demonstrating that the use of W2V provides better performance than the models used in the literature for tone identification on low-resource languages.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1145/3690384
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.