Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

May 02, 2024

Zongyang Du, Junchen Lu, Kun Zhou, Lakshmish Kaushik, Berrak Sisman

Figure 1 for Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

Figure 2 for Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

Figure 3 for Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

Figure 4 for Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

Share this with someone who'll enjoy it:

Abstract:Expressive voice conversion (VC) conducts speaker identity conversion for emotional speakers by jointly converting speaker identity and emotional style. Emotional style modeling for arbitrary speakers in expressive VC has not been extensively explored. Previous approaches have relied on vocoders for speech reconstruction, which makes speech quality heavily dependent on the performance of vocoders. A major challenge of expressive VC lies in emotion prosody modeling. To address these challenges, this paper proposes a fully end-to-end expressive VC framework based on a conditional denoising diffusion probabilistic model (DDPM). We utilize speech units derived from self-supervised speech models as content conditioning, along with deep features extracted from speech emotion recognition and speaker verification systems to model emotional style and speaker identity. Objective and subjective evaluations show the effectiveness of our framework. Codes and samples are publicly available.

* Accepted by Speaker Odyssey 2024

View paper on

Share this with someone who'll enjoy it:

Title:Converting Anyone's Voice: End-to-End Expressive Voice Conversion with a Conditional Diffusion Model

Paper and Code