Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Jan "Honza" Černocký

Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units

Jul 05, 2024

Bolaji Yusuf, Jan "Honza" Černocký, Murat Saraçlar

Figure 1 for Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units

Figure 2 for Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units

Figure 3 for Pretraining End-to-End Keyword Search with Automatically Discovered Acoustic Units

Abstract:End-to-end (E2E) keyword search (KWS) has emerged as an alternative and complimentary approach to conventional keyword search which depends on the output of automatic speech recognition (ASR) systems. While E2E methods greatly simplify the KWS pipeline, they generally have worse performance than their ASR-based counterparts, which can benefit from pretraining with untranscribed data. In this work, we propose a method for pretraining E2E KWS systems with untranscribed data, which involves using acoustic unit discovery (AUD) to obtain discrete units for untranscribed data and then learning to locate sequences of such units in the speech. We conduct experiments across languages and AUD systems: we show that finetuning such a model significantly outperforms a model trained from scratch, and the performance improvements are generally correlated with the quality of the AUD system used for pretraining.

* Interspeech 2024. KWS code at: https://github.com/bolajiy/golden-retriever; AUD code at https://github.com/beer-asr/beer/tree/master/recipes/hshmm

Via

Access Paper or Ask Questions