Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

Oct 30, 2020

Tsun-An Hsieh, Cheng Yu, Szu-Wei Fu, Xugang Lu, Yu Tsao

Figure 1 for Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

Figure 2 for Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

Figure 3 for Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

Figure 4 for Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

Share this with someone who'll enjoy it:

Abstract:Speech enhancement (SE) aims to improve speech quality and intelligibility, which are both related to a smooth transition in speech segments that may carry linguistic information, e.g. phones and syllables. In this study, we took phonetic characteristics into account in the SE training process. Hence, we designed a phone-fortified perceptual (PFP) loss, and the training of our SE model was guided by PFP loss. In PFP loss, phonetic characteristics are extracted by wav2vec, an unsupervised learning model based on the contrastive predictive coding (CPC) criterion. Different from previous deep-feature-based approaches, the proposed approach explicitly uses the phonetic information in the deep feature extraction process to guide the SE model training. To test the proposed approach, we first confirmed that the wav2vec representations carried clear phonetic information using a t-distributed stochastic neighbor embedding (t-SNE) analysis. Next, we observed that the proposed PFP loss was more strongly correlated with the perceptual evaluation metrics than point-wise and signal-level losses, thus achieving higher scores for standardized quality and intelligibility evaluation metrics in the Voice Bank-DEMAND dataset.

View paper on

Share this with someone who'll enjoy it:

Title:Improving Perceptual Quality by Phone-Fortified Perceptual Loss for Speech Enhancement

Paper and Code