Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

David van Leeuwen

Multi-Graph Decoding for Code-Switching ASR

Jun 28, 2019

Emre Yılmaz, Samuel Cohen, Xianghu Yue, David van Leeuwen, Haizhou Li

Figure 1 for Multi-Graph Decoding for Code-Switching ASR

Figure 2 for Multi-Graph Decoding for Code-Switching ASR

Figure 3 for Multi-Graph Decoding for Code-Switching ASR

Abstract:In the FAME! Project, a code-switching (CS) automatic speech recognition (ASR) system for Frisian-Dutch speech is developed that can accurately transcribe the local broadcaster's bilingual archives with CS speech. This archive contains recordings with monolingual Frisian and Dutch speech segments as well as Frisian-Dutch CS speech, hence the recognition performance on monolingual segments is also vital for accurate transcriptions. In this work, we propose a multi-graph decoding and rescoring strategy using bilingual and monolingual graphs together with a unified acoustic model for CS ASR. The proposed decoding scheme gives the freedom to design and employ alternative search spaces for each (monolingual or bilingual) recognition task and enables the effective use of monolingual resources of the high-resourced mixed language in low-resourced CS scenarios. In our scenario, Dutch is the high-resourced and Frisian is the low-resourced language. We therefore use additional monolingual Dutch text resources to improve the Dutch language model (LM) and compare the performance of single- and multi-graph CS ASR systems on Dutch segments using larger Dutch LMs. The ASR results show that the proposed approach outperforms baseline single-graph CS ASR systems, providing better performance on the monolingual Dutch segments without any accuracy loss on monolingual Frisian and code-mixed segments.

* Accepted for publication at Interspeech 2019

Via

Access Paper or Ask Questions

A comparison of linear and non-linear calibrations for speaker recognition

Apr 09, 2014

Niko Brümmer, Albert Swart, David van Leeuwen

Figure 1 for A comparison of linear and non-linear calibrations for speaker recognition

Figure 2 for A comparison of linear and non-linear calibrations for speaker recognition

Figure 3 for A comparison of linear and non-linear calibrations for speaker recognition

Figure 4 for A comparison of linear and non-linear calibrations for speaker recognition

Abstract:In recent work on both generative and discriminative score to log-likelihood-ratio calibration, it was shown that linear transforms give good accuracy only for a limited range of operating points. Moreover, these methods required tailoring of the calibration training objective functions in order to target the desired region of best accuracy. Here, we generalize the linear recipes to non-linear ones. We experiment with a non-linear, non-parametric, discriminative PAV solution, as well as parametric, generative, maximum-likelihood solutions that use Gaussian, Student's T and normal-inverse-Gaussian score distributions. Experiments on NIST SRE'12 scores suggest that the non-linear methods provide wider ranges of optimal accuracy and can be trained without having to resort to objective function tailoring.

* accepted for Odyssey 2014: The Speaker and Language Recognition Workshop

Via

Access Paper or Ask Questions