Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Transformer-based Models of Text Normalization for Speech Applications

Feb 01, 2022

Jae Hun Ro, Felix Stahlberg, Ke Wu, Shankar Kumar

Figure 1 for Transformer-based Models of Text Normalization for Speech Applications

Figure 2 for Transformer-based Models of Text Normalization for Speech Applications

Figure 3 for Transformer-based Models of Text Normalization for Speech Applications

Figure 4 for Transformer-based Models of Text Normalization for Speech Applications

Share this with someone who'll enjoy it:

Abstract:Text normalization, or the process of transforming text into a consistent, canonical form, is crucial for speech applications such as text-to-speech synthesis (TTS). In TTS, the system must decide whether to verbalize "1995" as "nineteen ninety five" in "born in 1995" or as "one thousand nine hundred ninety five" in "page 1995". We present an experimental comparison of various Transformer-based sequence-to-sequence (seq2seq) models of text normalization for speech and evaluate them on a variety of datasets of written text aligned to its normalized spoken form. These models include variants of the 2-stage RNN-based tagging/seq2seq architecture introduced by Zhang et al. (2019), where we replace the RNN with a Transformer in one or more stages, as well as vanilla Transformers that output string representations of edit sequences. Of our approaches, using Transformers for sentence context encoding within the 2-stage model proved most effective, with the fine-tuned BERT encoder yielding the best performance.

View paper on

Share this with someone who'll enjoy it:

Title:Transformer-based Models of Text Normalization for Speech Applications

Paper and Code