Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Elaine Sui

Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Jan 06, 2025

Yuhui Zhang, Yuchang Su, Yiming Liu, Xiaohan Wang, James Burgess, Elaine Sui, Chenyu Wang, Josiah Aklilu, Alejandro Lozano, Anjiang Wei(+2 more)

Figure 1 for Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Figure 2 for Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Figure 3 for Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Figure 4 for Automated Generation of Challenging Multiple-Choice Questions for Vision Language Model Evaluation

Abstract:The rapid development of vision language models (VLMs) demands rigorous and reliable evaluation. However, current visual question answering (VQA) benchmarks often depend on open-ended questions, making accurate evaluation difficult due to the variability in natural language responses. To address this, we introduce AutoConverter, an agentic framework that automatically converts these open-ended questions into multiple-choice format, enabling objective evaluation while reducing the costly question creation process. Our experiments demonstrate that AutoConverter can generate correct and challenging multiple-choice questions, with VLMs demonstrating consistently similar or lower accuracy on these questions compared to human-created ones. Using AutoConverter, we construct VMCBench, a benchmark created by transforming 20 existing VQA datasets into a unified multiple-choice format, totaling 9,018 questions. We comprehensively evaluate 33 state-of-the-art VLMs on VMCBench, setting a new standard for scalable, consistent, and reproducible VLM evaluation.

* Project page: https://yuhui-zh15.github.io/AutoConverter-Website/

Via

Access Paper or Ask Questions

Just Shift It: Test-Time Prototype Shifting for Zero-Shot Generalization with Vision-Language Models

Mar 19, 2024

Elaine Sui, Xiaohan Wang, Serena Yeung-Levy

Abstract:Advancements in vision-language models (VLMs) have propelled the field of computer vision, particularly in the zero-shot learning setting. Despite their promise, the effectiveness of these models often diminishes due to domain shifts in test environments. To address this, we introduce the Test-Time Prototype Shifting (TPS) framework, a pioneering approach designed to adapt VLMs to test datasets using unlabeled test inputs. Our method is based on the notion of modulating per-class prototypes in the shared embedding space. By pre-computing and caching prototypes generated with the pre-trained text encoder, TPS not only facilitates optimization-free prototype reuse for subsequent predictions but also enables seamless integration with current advancements in prompt engineering. At test-time, TPS dynamically learns shift vectors for each prototype based solely on the given test sample, effectively bridging the domain gap and enhancing classification accuracy. A notable aspect of our framework is its significantly reduced memory and computational demands when compared to conventional text-prompt tuning methods. Extensive evaluations across 15 datasets involving natural distribution shifts and cross-dataset generalization demonstrate TPS's superior performance, achieving state-of-the-art results while reducing resource requirements.

Via

Access Paper or Ask Questions

Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Jan 16, 2024

Yuhui Zhang, Elaine Sui, Serena Yeung-Levy

Figure 1 for Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Figure 2 for Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Figure 3 for Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Figure 4 for Connect, Collapse, Corrupt: Learning Cross-Modal Tasks with Uni-Modal Data

Abstract:Building cross-modal applications is challenging due to limited paired multi-modal data. Recent works have shown that leveraging a pre-trained multi-modal contrastive representation space enables cross-modal tasks to be learned from uni-modal data. This is based on the assumption that contrastive optimization makes embeddings from different modalities interchangeable. However, this assumption is under-explored due to the poorly understood geometry of the multi-modal contrastive space, where a modality gap exists. In our study, we provide a theoretical explanation of this space's geometry and introduce a three-step method, $C^3$ (Connect, Collapse, Corrupt), to bridge the modality gap, enhancing the interchangeability of embeddings. Our $C^3$ method significantly improves cross-modal learning from uni-modal data, achieving state-of-the-art results on zero-shot image / audio / video captioning and text-to-image generation.

* Published at ICLR 2024

Via

Access Paper or Ask Questions