article
The main goal of this paper is the transfer of expressivity from a reference speech to a synthesized speech. The presented approach conditions the transformer network text to speech synthesis system for transferring prosody from a reference audio utterance to a spoken text. We propose to extend the transformer TTS architecture with a prosody encoder and a multi-head cross-attention block used for fusion and alignment between text and expressivity content. We compare our model to the baseline models of transfer of expressivity namely Global Style Token (GST), Variational AutoEncoder (VAE), and Fine- Grained style control in transformer TTS (FGT). The proposed model shows good prosody transfer and enhances the expressivity and quality of the generated speech as illustrated by listening tests and correlated subjective and objective metrics for evaluating prosody transfer task and naturalness of speech.
This page summarises published work. The authoritative version sits with the publisher.
DOI: 10.1109/eeite61750.2024.10654450
Is something wrong with this record? Report it or request removal.
Discussion
Have you built on this work, tried to replicate it, or seen it applied in practice? Share what you know. Verified researchers and MARATTO™ domain experts can open a discussion, and any member can reply. Contributions are reviewed before they appear.
No discussion yet. Open the first thread.
New to MARATTO™? Create a free account.