Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Grounded Image Text Matching with Mismatched Relation Reasoning

Aug 04, 2023

Yu Wu, Yana Wei, Haozhe Wang, Yongfei Liu, Sibei Yang, Xuming He

Figure 1 for Grounded Image Text Matching with Mismatched Relation Reasoning

Figure 2 for Grounded Image Text Matching with Mismatched Relation Reasoning

Figure 3 for Grounded Image Text Matching with Mismatched Relation Reasoning

Figure 4 for Grounded Image Text Matching with Mismatched Relation Reasoning

Share this with someone who'll enjoy it:

Abstract:This paper introduces Grounded Image Text Matching with Mismatched Relation (GITM-MR), a novel visual-linguistic joint task that evaluates the relation understanding capabilities of transformer-based pre-trained models. GITM-MR requires a model to first determine if an expression describes an image, then localize referred objects or ground the mismatched parts of the text. We provide a benchmark for evaluating pre-trained models on this task, with a focus on the challenging settings of limited data and out-of-distribution sentence lengths. Our evaluation demonstrates that pre-trained models lack data efficiency and length generalization ability. To address this, we propose the Relation-sensitive Correspondence Reasoning Network (RCRN), which incorporates relation-aware reasoning via bi-directional message propagation guided by language structure. RCRN can be interpreted as a modular program and delivers strong performance in both length generalization and data efficiency.

View paper on

Share this with someone who'll enjoy it:

Title:Grounded Image Text Matching with Mismatched Relation Reasoning

Paper and Code