MARATTO

preprint · Cognitive Computation

A Multilingual Training Strategy for Low-Resource Text-to-Speech

2026Open accessMohammed V University

Abstract

Recent speech technologies have led to the production of high quality synthesised speech due to recent advances in neural text-to-speech (TTS). However, such TTS models depend on extensive amounts of data that can be costly to produce and are hardly scalable to all existing languages, especially since little attention is given to low-resource languages. With techniques such as knowledge transfer, the burden of creating datasets can be alleviated. In this paper, we therefore investigate two aspects; firstly, whether data from social media can be used for a small TTS dataset construction, and secondly whether cross-lingual transfer learning (TL) for a low-resource language can work with this type of data. In this aspect, we specifically assess to what extent multilingual modeling can be leveraged as an alternative to training on monolingual corpora. To do so, we explore how data from foreign languages may be selected and pooled to train a TTS model for a target low-resource language. Our findings show that multilingual pre-training with informed source language selection outperforms monolingual pre-training in enhancing both the intelligibility and naturalness of the generated speech.

Research topics

  • Speech and dialogue systems
  • Natural Language Processing Techniques

Read the original research

This page summarises published work. The authoritative version sits with the publisher.

DOI: 10.1007/s12559-026-10598-3

Is something wrong with this record? Report it or request removal.

Discussion

Discuss this research

Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.

No discussion yet. Open the first thread.