Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering

Aug 28, 2021

Gouthaman KV, Anurag Mittal

Figure 1 for On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering

Figure 2 for On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering

Figure 3 for On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering

Figure 4 for On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering

Share this with someone who'll enjoy it:

Abstract:Generalizing beyond the experiences has a significant role in developing practical AI systems. It has been shown that current Visual Question Answering (VQA) models are over-dependent on the language-priors (spurious correlations between question-types and their most frequent answers) from the train set and pose poor performance on Out-of-Distribution (OOD) test sets. This conduct limits their generalizability and restricts them from being utilized in real-world situations. This paper shows that the sequence model architecture used in the question-encoder has a significant role in the generalizability of VQA models. To demonstrate this, we performed a detailed analysis of various existing RNN-based and Transformer-based question-encoders, and along, we proposed a novel Graph attention network (GAT)-based question-encoder. Our study found that a better choice of sequence model in the question-encoder improves the generalizability of VQA models even without using any additional relatively complex bias-mitigation approaches.

View paper on

Share this with someone who'll enjoy it:

Title:On the Significance of Question Encoder Sequence Model in the Out-of-Distribution Performance in Visual Question Answering

Paper and Code