Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Are Multimodal Models Robust to Image and Text Perturbations?

Dec 15, 2022

Jielin Qiu, Yi Zhu, Xingjian Shi, Florian Wenzel, Zhiqiang Tang, Ding Zhao, Bo Li, Mu Li

Figure 1 for Are Multimodal Models Robust to Image and Text Perturbations?

Figure 2 for Are Multimodal Models Robust to Image and Text Perturbations?

Figure 3 for Are Multimodal Models Robust to Image and Text Perturbations?

Figure 4 for Are Multimodal Models Robust to Image and Text Perturbations?

Share this with someone who'll enjoy it:

Abstract:Multimodal image-text models have shown remarkable performance in the past few years. However, evaluating their robustness against distribution shifts is crucial before adopting them in real-world applications. In this paper, we investigate the robustness of 9 popular open-sourced image-text models under common perturbations on five tasks (image-text retrieval, visual reasoning, visual entailment, image captioning, and text-to-image generation). In particular, we propose several new multimodal robustness benchmarks by applying 17 image perturbation and 16 text perturbation techniques on top of existing datasets. We observe that multimodal models are not robust to image and text perturbations, especially to image perturbations. Among the tested perturbation methods, character-level perturbations constitute the most severe distribution shift for text, and zoom blur is the most severe shift for image data. We also introduce two new robustness metrics (MMI and MOR) for proper evaluations of multimodal models. We hope our extensive study sheds light on new directions for the development of robust multimodal models.

* The project webpage is at: https://mmrobustness.github.io/

View paper on

Share this with someone who'll enjoy it:

Title:Are Multimodal Models Robust to Image and Text Perturbations?

Paper and Code