Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Text-Aware Dual Routing Network for Visual Question Answering

Nov 17, 2022

Luoqian Jiang, Yifan He, Jian Chen

Figure 1 for Text-Aware Dual Routing Network for Visual Question Answering

Figure 2 for Text-Aware Dual Routing Network for Visual Question Answering

Figure 3 for Text-Aware Dual Routing Network for Visual Question Answering

Figure 4 for Text-Aware Dual Routing Network for Visual Question Answering

Share this with someone who'll enjoy it:

Abstract:Visual question answering (VQA) is a challenging task to provide an accurate natural language answer given an image and a natural language question about the image. It involves multi-modal learning, i.e., computer vision (CV) and natural language processing (NLP), as well as flexible answer prediction for free-form and open-ended answers. Existing approaches often fail in cases that require reading and understanding text in images to answer questions. In practice, they cannot effectively handle the answer sequence derived from text tokens because the visual features are not text-oriented. To address the above issues, we propose a Text-Aware Dual Routing Network (TDR) which simultaneously handles the VQA cases with and without understanding text information in the input images. Specifically, we build a two-branch answer prediction network that contains a specific branch for each case and further develop a dual routing scheme to dynamically determine which branch should be chosen. In the branch that involves text understanding, we incorporate the Optical Character Recognition (OCR) features into the model to help understand the text in the images. Extensive experiments on the VQA v2.0 dataset demonstrate that our proposed TDR outperforms existing methods, especially on the ''number'' related VQA questions.

View paper on

Share this with someone who'll enjoy it:

Title:Text-Aware Dual Routing Network for Visual Question Answering

Paper and Code