Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:On-demand compute reduction with stochastic wav2vec 2.0

Apr 25, 2022

Apoorv Vyas, Wei-Ning Hsu, Michael Auli, Alexei Baevski

Figure 1 for On-demand compute reduction with stochastic wav2vec 2.0

Figure 2 for On-demand compute reduction with stochastic wav2vec 2.0

Figure 3 for On-demand compute reduction with stochastic wav2vec 2.0

Figure 4 for On-demand compute reduction with stochastic wav2vec 2.0

Share this with someone who'll enjoy it:

Abstract:Squeeze and Efficient Wav2vec (SEW) is a recently proposed architecture that squeezes the input to the transformer encoder for compute efficient pre-training and inference with wav2vec 2.0 (W2V2) models. In this work, we propose stochastic compression for on-demand compute reduction for W2V2 models. As opposed to using a fixed squeeze factor, we sample it uniformly during training. We further introduce query and key-value pooling mechanisms that can be applied to each transformer layer for further compression. Our results for models pre-trained on 960h Librispeech dataset and fine-tuned on 10h of transcribed data show that using the same stochastic model, we get a smooth trade-off between word error rate (WER) and inference time with only marginal WER degradation compared to the W2V2 and SEW models trained for a specific setting. We further show that we can fine-tune the same stochastically pre-trained model to a specific configuration to recover the WER difference resulting in significant computational savings on pre-training models from scratch.

* submitted to Interspeech, 2022

View paper on

Share this with someone who'll enjoy it:

Title:On-demand compute reduction with stochastic wav2vec 2.0

Paper and Code