Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Nov 21, 2024

Yuke Zhu, Chi Xie, Shuang Liang, Bo Zheng, Sheng Guo

Figure 1 for FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Figure 2 for FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Figure 3 for FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Figure 4 for FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Share this with someone who'll enjoy it:

Abstract:Recent advances on Multi-modal Large Language Models have demonstrated that high-resolution image input is crucial for model capabilities, especially for fine-grained tasks. However, high-resolution images lead to a quadratic increase in the number of visual tokens input into LLMs, resulting in significant computational costs. Current work develop visual token compression methods to achieve efficiency improvements, often at the expense of performance. We argue that removing visual redundancy can simultaneously improve both efficiency and performance. We build a coarse-to-fine visual token compression method, with a vision-guided sampler for compressing redundant regions with low information density, and a text-guided sampler for selecting visual tokens that are strongly correlated with the user instructions.With these two modules, the proposed FocusLLaVA achieves improvements in both efficiency and performance. We validate the effectiveness of our approach on a wide range of evaluation datasets.

View paper on

Share this with someone who'll enjoy it:

Title:FocusLLaVA: A Coarse-to-Fine Approach for Efficient and Effective Visual Token Compression

Paper and Code