Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

Jun 25, 2022

Yihan Wu, Xi Wang, Shaofei Zhang, Lei He, Ruihua Song, Jian-Yun Nie

Figure 1 for Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

Figure 2 for Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

Figure 3 for Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

Figure 4 for Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

Share this with someone who'll enjoy it:

Abstract:Expressive speech synthesis, like audiobook synthesis, is still challenging for style representation learning and prediction. Deriving from reference audio or predicting style tags from text requires a huge amount of labeled data, which is costly to acquire and difficult to define and annotate accurately. In this paper, we propose a novel framework for learning style representation from abundant plain text in a self-supervised manner. It leverages an emotion lexicon and uses contrastive learning and deep clustering. We further integrate the style representation as a conditioned embedding in a multi-style Transformer TTS. Comparing with multi-style TTS by predicting style tags trained on the same dataset but with human annotations, our method achieves improved results according to subjective evaluations on both in-domain and out-of-domain test sets in audiobook speech. Moreover, with implicit context-aware style representation, the emotion transition of synthesized audio in a long paragraph appears more natural. The audio samples are available on the demo web.

* Accepted by Interspeech 2022

View paper on

Share this with someone who'll enjoy it:

Title:Self-supervised Context-aware Style Representation for Expressive Speech Synthesis

Paper and Code