Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition

Jun 23, 2020

Dongwei Jiang, Wubo Li, Ruixiong Zhang, Miao Cao, Ne Luo, Yang Han, Wei Zou, Xiangang Li

Figure 1 for A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition

Figure 2 for A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition

Figure 3 for A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition

Figure 4 for A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition

Share this with someone who'll enjoy it:

Abstract:Building a good speech recognition system usually requires large amounts of transcribed data, which is expensive to collect. To tackle this problem, many unsupervised pre-training methods have been proposed. Among these methods, Masked Predictive Coding achieved significant improvements on various speech recognition datasets with BERT-like Masked Reconstruction loss and Transformer backbone. However, many aspects of MPC have not been fully investigated. In this paper, we conduct a further study on MPC and focus on three important aspects: the effect of pre-training data speaking style, its extension on streaming model, and how to better transfer learned knowledge from pre-training stage to downstream tasks. Experiments reveled that pre-training data with matching speaking style is more useful on downstream recognition tasks. A unified training objective with APC and MPC provided 8.46% relative error reduction on streaming model trained on HKUST. Also, the combination of target data adaption and layer-wise discriminative training helped the knowledge transfer of MPC, which achieved 3.99% relative error reduction on AISHELL over a strong baseline.

View paper on

OpenReview

Share this with someone who'll enjoy it:

Title:A Further Study of Unsupervised Pre-training for Transformer Based Speech Recognition

Paper and Code