Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Miguel del Rio

Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

Sep 04, 2024

Annika Heuser, Tyler Kendall, Miguel del Rio, Quinten McNamara, Nishchal Bhandari, Corey Miller, Migüel Jetté

Figure 1 for Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

Figure 2 for Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

Figure 3 for Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

Figure 4 for Quantification of stylistic differences in human- and ASR-produced transcripts of African American English

Abstract:Common measures of accuracy used to assess the performance of automatic speech recognition (ASR) systems, as well as human transcribers, conflate multiple sources of error. Stylistic differences, such as verbatim vs non-verbatim, can play a significant role in ASR performance evaluation when differences exist between training and test datasets. The problem is compounded for speech from underrepresented varieties, where the speech to orthography mapping is not as standardized. We categorize the kinds of stylistic differences between 6 transcription versions, 4 human- and 2 ASR-produced, of 10 hours of African American English (AAE) speech. Focusing on verbatim features and AAE morphosyntactic features, we investigate the interactions of these categories with how well transcripts can be compared via word error rate (WER). The results, and overall analysis, help clarify how ASR outputs are a function of the decisions made by the training data's human transcribers.

* Published in Interspeech 2024 Proceedings, 5 pages excluding references, 5 figures

Via

Access Paper or Ask Questions