Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

Apr 26, 2024

Yuhang Huang, Zihan Wu, Chongyang Gao, Jiawei Peng, Xu Yang

Figure 1 for Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

Figure 2 for Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

Figure 3 for Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

Figure 4 for Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

Share this with someone who'll enjoy it:

Abstract:Large Vision-Language Models (LVLMs) are gaining traction for their remarkable ability to process and integrate visual and textual data. Despite their popularity, the capacity of LVLMs to generate precise, fine-grained textual descriptions has not been fully explored. This study addresses this gap by focusing on \textit{distinctiveness} and \textit{fidelity}, assessing how models like Open-Flamingo, IDEFICS, and MiniGPT-4 can distinguish between similar objects and accurately describe visual features. We proposed the Textual Retrieval-Augmented Classification (TRAC) framework, which, by leveraging its generative capabilities, allows us to delve deeper into analyzing fine-grained visual description generation. This research provides valuable insights into the generation quality of LVLMs, enhancing the understanding of multimodal language models. Notably, MiniGPT-4 stands out for its better ability to generate fine-grained descriptions, outperforming the other two models in this aspect. The code is provided at \url{https://anonymous.4open.science/r/Explore_FGVDs-E277}.

* 11 pages, 9 figures, 6 tables. For associated code, see https://anonymous.4open.science/r/Explore_FGVDs-E277

View paper on

Share this with someone who'll enjoy it:

Title:Exploring the Distinctiveness and Fidelity of the Descriptions Generated by Large Vision-Language Models

Paper and Code