Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering

Jan 11, 2022

Ankur Sikarwar, Gabriel Kreiman

Figure 1 for On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering

Figure 2 for On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering

Figure 3 for On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering

Figure 4 for On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering

Share this with someone who'll enjoy it:

Abstract:In recent years, multi-modal transformers have shown significant progress in Vision-Language tasks, such as Visual Question Answering (VQA), outperforming previous architectures by a considerable margin. This improvement in VQA is often attributed to the rich interactions between vision and language streams. In this work, we investigate the efficacy of co-attention transformer layers in helping the network focus on relevant regions while answering the question. We generate visual attention maps using the question-conditioned image attention scores in these co-attention layers. We evaluate the effect of the following critical components on visual attention of a state-of-the-art VQA model: (i) number of object region proposals, (ii) question part of speech (POS) tags, (iii) question semantics, (iv) number of co-attention layers, and (v) answer accuracy. We compare the neural network attention maps against human attention maps both qualitatively and quantitatively. Our findings indicate that co-attention transformer modules are crucial in attending to relevant regions of the image given a question. Importantly, we observe that the semantic meaning of the question is not what drives visual attention, but specific keywords in the question do. Our work sheds light on the function and interpretation of co-attention transformer layers, highlights gaps in current networks, and can guide the development of future VQA models and networks that simultaneously process visual and language streams.

View paper on

Share this with someone who'll enjoy it:

Title:On the Efficacy of Co-Attention Transformer Layers in Visual Question Answering

Paper and Code