Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Nagendra Kumar Goel

Comparison of SVD and factorized TDNN approaches for speech to text

Oct 13, 2021

Jeffrey Josanne Michael, Nagendra Kumar Goel, Navneeth K, Jonas Robertson, Shravan Mishra

Figure 1 for Comparison of SVD and factorized TDNN approaches for speech to text

Figure 2 for Comparison of SVD and factorized TDNN approaches for speech to text

Figure 3 for Comparison of SVD and factorized TDNN approaches for speech to text

Figure 4 for Comparison of SVD and factorized TDNN approaches for speech to text

Abstract:This work concentrates on reducing the RTF and word error rate of a hybrid HMM-DNN. Our baseline system uses an architecture with TDNN and LSTM layers. We find this architecture particularly useful for lightly reverberated environments. However, these models tend to demand more computation than is desirable. In this work, we explore alternate architectures employing singular value decomposition (SVD) is applied to the TDNN layers to reduce the RTF, as well as to the affine transforms of every LSTM cell. We compare this approach with specifying bottleneck layers similar to those introduced by SVD before training. Additionally, we reduced the search space of the decoding graph to make it a better fit to operate in real-time applications. We report -61.57% relative reduction in RTF and almost 1% relative decrease in WER for our architecture trained on Fisher data along with reverberated versions of this dataset in order to match one of our target test distributions.

* 4 pages, 1 figure, 3 tables

Via

Access Paper or Ask Questions