Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Abhishek Chaudhary

Semantically Distributed Robust Optimization for Vision-and-Language Inference

Oct 14, 2021

Tejas Gokhale, Abhishek Chaudhary, Pratyay Banerjee, Chitta Baral, Yezhou Yang

Figure 1 for Semantically Distributed Robust Optimization for Vision-and-Language Inference

Figure 2 for Semantically Distributed Robust Optimization for Vision-and-Language Inference

Figure 3 for Semantically Distributed Robust Optimization for Vision-and-Language Inference

Figure 4 for Semantically Distributed Robust Optimization for Vision-and-Language Inference

Abstract:Analysis of vision-and-language models has revealed their brittleness under linguistic phenomena such as paraphrasing, negation, textual entailment, and word substitutions with synonyms or antonyms. While data augmentation techniques have been designed to mitigate against these failure modes, methods that can integrate this knowledge into the training pipeline remain under-explored. In this paper, we present \textbf{SDRO}, a model-agnostic method that utilizes a set linguistic transformations in a distributed robust optimization setting, along with an ensembling technique to leverage these transformations during inference. Experiments on benchmark datasets with images (NLVR$^2$) and video (VIOLIN) demonstrate performance improvements as well as robustness to adversarial attacks. Experiments on binary VQA explore the generalizability of this method to other V\&L tasks.

* preprint; code available at https://github.com/ASU-APG/VLI_SDRO

Via

Access Paper or Ask Questions