Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Aug 29, 2024

Zaiwei Zhang, Gregory P. Meyer, Zhichao Lu, Ashish Shrivastava, Avinash Ravichandran, Eric M. Wolff

Figure 1 for VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Figure 2 for VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Figure 3 for VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Figure 4 for VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Share this with someone who'll enjoy it:

Abstract:For visual recognition, knowledge distillation typically involves transferring knowledge from a large, well-trained teacher model to a smaller student model. In this paper, we introduce an effective method to distill knowledge from an off-the-shelf vision-language model (VLM), demonstrating that it provides novel supervision in addition to those from a conventional vision-only teacher model. Our key technical contribution is the development of a framework that generates novel text supervision and distills free-form text into a vision encoder. We showcase the effectiveness of our approach, termed VLM-KD, across various benchmark datasets, showing that it surpasses several state-of-the-art long-tail visual classifiers. To our knowledge, this work is the first to utilize knowledge distillation with text supervision generated by an off-the-shelf VLM and apply it to vanilla randomly initialized vision encoders.

View paper on

Share this with someone who'll enjoy it:

Title:VLM-KD: Knowledge Distillation from VLM for Long-Tail Visual Recognition

Paper and Code