Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Boosting Audio-visual Zero-shot Learning with Large Language Models

Nov 21, 2023

Haoxing Chen, Yaohui Li, Yan Hong, Zizheng Huang, Zhuoer Xu, Zhangxuan Gu, Jun Lan, Huijia Zhu, Weiqiang Wang

Figure 1 for Boosting Audio-visual Zero-shot Learning with Large Language Models

Figure 2 for Boosting Audio-visual Zero-shot Learning with Large Language Models

Figure 3 for Boosting Audio-visual Zero-shot Learning with Large Language Models

Figure 4 for Boosting Audio-visual Zero-shot Learning with Large Language Models

Share this with someone who'll enjoy it:

Abstract:Audio-visual zero-shot learning aims to recognize unseen categories based on paired audio-visual sequences. Recent methods mainly focus on learning aligned and discriminative multi-modal features to boost generalization towards unseen categories. However, these approaches ignore the obscure action concepts in category names and may inevitably introduce complex network structures with difficult training objectives. In this paper, we propose a simple yet effective framework named Knowledge-aware Distribution Adaptation (KDA) to help the model better grasp the novel action contents with an external knowledge base. Specifically, we first propose using large language models to generate rich descriptions from category names, which leads to a better understanding of unseen categories. Additionally, we propose a distribution alignment loss as well as a knowledge-aware adaptive margin loss to further improve the generalization ability towards unseen categories. Extensive experimental results demonstrate that our proposed KDA can outperform state-of-the-art methods on three popular audio-visual zero-shot learning datasets. Our code will be avaliable at \url{https://github.com/chenhaoxing/KDA}.

View paper on

Share this with someone who'll enjoy it:

Title:Boosting Audio-visual Zero-shot Learning with Large Language Models

Paper and Code