Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Apr 16, 2024

Yanze Li, Wenhua Zhang, Kai Chen, Yanxin Liu, Pengxiang Li, Ruiyuan Gao, Lanqing Hong, Meng Tian, Xinhai Zhao, Zhenguo Li(+3 more)

Figure 1 for Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Figure 2 for Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Figure 3 for Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Figure 4 for Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Share this with someone who'll enjoy it:

Abstract:Large Vision-Language Models (LVLMs), due to the remarkable visual reasoning ability to understand images and videos, have received widespread attention in the autonomous driving domain, which significantly advances the development of interpretable end-to-end autonomous driving. However, current evaluations of LVLMs primarily focus on the multi-faceted capabilities in common scenarios, lacking quantifiable and automated assessment in autonomous driving contexts, let alone severe road corner cases that even the state-of-the-art autonomous driving perception systems struggle to handle. In this paper, we propose CODA-LM, a novel vision-language benchmark for self-driving, which provides the first automatic and quantitative evaluation of LVLMs for interpretable autonomous driving including general perception, regional perception, and driving suggestions. CODA-LM utilizes the texts to describe the road images, exploiting powerful text-only large language models (LLMs) without image inputs to assess the capabilities of LVLMs in autonomous driving scenarios, which reveals stronger alignment with human preferences than LVLM judges. Experiments demonstrate that even the closed-sourced commercial LVLMs like GPT-4V cannot deal with road corner cases well, suggesting that we are still far from a strong LVLM-powered intelligent driving agent, and we hope our CODA-LM can become the catalyst to promote future development.

* Project Page: https://coda-dataset.github.io/coda-lm/

View paper on

Share this with someone who'll enjoy it:

Title:Automated Evaluation of Large Vision-Language Models on Self-driving Corner Cases

Paper and Code