Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Multilingual Turn-taking Prediction Using Voice Activity Projection

Mar 14, 2024

Koji Inoue, Bing'er Jiang, Erik Ekstedt, Tatsuya Kawahara, Gabriel Skantze

Figure 1 for Multilingual Turn-taking Prediction Using Voice Activity Projection

Figure 2 for Multilingual Turn-taking Prediction Using Voice Activity Projection

Figure 3 for Multilingual Turn-taking Prediction Using Voice Activity Projection

Figure 4 for Multilingual Turn-taking Prediction Using Voice Activity Projection

Share this with someone who'll enjoy it:

Abstract:This paper investigates the application of voice activity projection (VAP), a predictive turn-taking model for spoken dialogue, on multilingual data, encompassing English, Mandarin, and Japanese. The VAP model continuously predicts the upcoming voice activities of participants in dyadic dialogue, leveraging a cross-attention Transformer to capture the dynamic interplay between participants. The results show that a monolingual VAP model trained on one language does not make good predictions when applied to other languages. However, a multilingual model, trained on all three languages, demonstrates predictive performance on par with monolingual models across all languages. Further analyses show that the multilingual model has learned to discern the language of the input signal. We also analyze the sensitivity to pitch, a prosodic cue that is thought to be important for turn-taking. Finally, we compare two different audio encoders, contrastive predictive coding (CPC) pre-trained on English, with a recent model based on multilingual wav2vec 2.0 (MMS).

* This paper has been accepted for presentation at The 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024) and represents the author's version of the work

View paper on

Share this with someone who'll enjoy it:

Title:Multilingual Turn-taking Prediction Using Voice Activity Projection

Paper and Code