Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Jul 13, 2022

Yookyung Shin, Younggun Lee, Suhee Jo, Yeongtae Hwang, Taesu Kim

Figure 1 for Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Figure 2 for Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Figure 3 for Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Figure 4 for Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Share this with someone who'll enjoy it:

Abstract:Expressive text-to-speech has shown improved performance in recent years. However, the style control of synthetic speech is often restricted to discrete emotion categories and requires training data recorded by the target speaker in the target style. In many practical situations, users may not have reference speech recorded in target emotion but still be interested in controlling speech style just by typing text description of desired emotional style. In this paper, we propose a text-based interface for emotional style control and cross-speaker style transfer in multi-speaker TTS. We propose the bi-modal style encoder which models the semantic relationship between text description embedding and speech style embedding with a pretrained language model. To further improve cross-speaker style transfer on disjoint, multi-style datasets, we propose the novel style loss. The experimental results show that our model can generate high-quality expressive speech even in unseen style.

* Accepted to Interspeech 2022

View paper on

Share this with someone who'll enjoy it:

Title:Text-driven Emotional Style Control and Cross-speaker Style Transfer in Neural TTS

Paper and Code