Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

May 18, 2023

Hao Shi, Kazuki Shimada, Masato Hirano, Takashi Shibuya, Yuichiro Koyama, Zhi Zhong, Shusuke Takahashi, Tatsuya Kawahara, Yuki Mitsufuji

Figure 1 for Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

Figure 2 for Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

Figure 3 for Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

Figure 4 for Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

Share this with someone who'll enjoy it:

Abstract:Diffusion-based speech enhancement (SE) has been investigated recently, but its decoding is very time-consuming. One solution is to initialize the decoding process with the enhanced feature estimated by a predictive SE system. However, this two-stage method ignores the complementarity between predictive and diffusion SE. In this paper, we propose a unified system that integrates these two SE modules. The system encodes both generative and predictive information, and then applies both generative and predictive decoders, whose outputs are fused. Specifically, the two SE modules are fused in the first and final diffusion steps: the first step fusion initializes the diffusion process with the predictive SE for improving the convergence, and the final step fusion combines the two complementary SE outputs to improve the SE performance. Experiments on the Voice-Bank dataset show that the diffusion score estimation can benefit from the predictive information and speed up the decoding.

View paper on

Share this with someone who'll enjoy it:

Title:Diffusion-Based Speech Enhancement with Joint Generative and Predictive Decoders

Paper and Code