Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Self-Supervised Speech Representations are More Phonetic than Semantic

Jun 12, 2024

Kwanghee Choi, Ankita Pasad, Tomohiko Nakamura, Satoru Fukayama, Karen Livescu, Shinji Watanabe

Figure 1 for Self-Supervised Speech Representations are More Phonetic than Semantic

Figure 2 for Self-Supervised Speech Representations are More Phonetic than Semantic

Figure 3 for Self-Supervised Speech Representations are More Phonetic than Semantic

Figure 4 for Self-Supervised Speech Representations are More Phonetic than Semantic

Share this with someone who'll enjoy it:

Abstract:Self-supervised speech models (S3Ms) have become an effective backbone for speech applications. Various analyses suggest that S3Ms encode linguistic properties. In this work, we seek a more fine-grained analysis of the word-level linguistic properties encoded in S3Ms. Specifically, we curate a novel dataset of near homophone (phonetically similar) and synonym (semantically similar) word pairs and measure the similarities between S3M word representation pairs. Our study reveals that S3M representations consistently and significantly exhibit more phonetic than semantic similarity. Further, we question whether widely used intent classification datasets such as Fluent Speech Commands and Snips Smartlights are adequate for measuring semantic abilities. Our simple baseline, using only the word identity, surpasses S3M-based models. This corroborates our findings and suggests that high scores on these datasets do not necessarily guarantee the presence of semantic content.

* Accepted to Interspeech 2024. Source code at https://github.com/juice500ml/phonetic_semantic_probing

View paper on

Share this with someone who'll enjoy it:

Title:Self-Supervised Speech Representations are More Phonetic than Semantic

Paper and Code