Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Jan 15, 2021

Violetta Shevchenko, Damien Teney, Anthony Dick, Anton van den Hengel

Figure 1 for Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Figure 2 for Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Figure 3 for Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Figure 4 for Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Share this with someone who'll enjoy it:

Abstract:The limits of applicability of vision-and-language models are defined by the coverage of their training data. Tasks like vision question answering (VQA) often require commonsense and factual information beyond what can be learned from task-specific datasets. This paper investigates the injection of knowledge from general-purpose knowledge bases (KBs) into vision-and-language transformers. We use an auxiliary training objective that encourages the learned representations to align with graph embeddings of matching entities in a KB. We empirically study the relevance of various KBs to multiple tasks and benchmarks. The technique brings clear benefits to knowledge-demanding question answering tasks (OK-VQA, FVQA) by capturing semantic and relational knowledge absent from existing models. More surprisingly, the technique also benefits visual reasoning tasks (NLVR2, SNLI-VE). We perform probing experiments and show that the injection of additional knowledge regularizes the space of embeddings, which improves the representation of lexical and semantic similarities. The technique is model-agnostic and can expand the applicability of any vision-and-language transformer with minimal computational overhead.

View paper on

Share this with someone who'll enjoy it:

Title:Reasoning over Vision and Language: Exploring the Benefits of Supplemental Knowledge

Paper and Code