Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Sep 29, 2021

Rui Li, Dong Pu, Minnie Huang, Bill Huang

Figure 1 for Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Figure 2 for Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Figure 3 for Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Figure 4 for Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Share this with someone who'll enjoy it:

Abstract:One-shot voice cloning aims to transform speaker voice and speaking style in speech synthesized from a text-to-speech (TTS) system, where only a shot recording from the target reference speech can be used. Out-of-domain transfer is still a challenging task, and one important aspect that impacts the accuracy and similarity of synthetic speech is the conditional representations carrying speaker or style cues extracted from the limited references. In this paper, we present a novel one-shot voice cloning algorithm called Unet-TTS that has good generalization ability for unseen speakers and styles. Based on a skip-connected U-net structure, the new model can efficiently discover speaker-level and utterance-level spectral feature details from the reference audio, enabling accurate inference of complex acoustic characteristics as well as imitation of speaking styles into the synthetic speech. According to both subjective and objective evaluations of similarity, the new model outperforms both speaker embedding and unsupervised style modeling (GST) approaches on an unseen emotional corpus.

* 6 pages, 5 figures, Submitted to IEEE ICASSP 2022

View paper on

Share this with someone who'll enjoy it:

Title:Unet-TTS: Improving Unseen Speaker and Style Transfer in One-shot Voice Cloning

Paper and Code