Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:R$^3$Net:Relation-embedded Representation Reconstruction Network for Change Captioning

Oct 20, 2021

Yunbin Tu, Liang Li, Chenggang Yan, Shengxiang Gao, Zhengtao Yu

Figure 1 for R$^3$Net:Relation-embedded Representation Reconstruction Network for Change Captioning

Figure 2 for R$^3$Net:Relation-embedded Representation Reconstruction Network for Change Captioning

Figure 3 for R$^3$Net:Relation-embedded Representation Reconstruction Network for Change Captioning

Figure 4 for R$^3$Net:Relation-embedded Representation Reconstruction Network for Change Captioning

Share this with someone who'll enjoy it:

Abstract:Change captioning is to use a natural language sentence to describe the fine-grained disagreement between two similar images. Viewpoint change is the most typical distractor in this task, because it changes the scale and location of the objects and overwhelms the representation of real change. In this paper, we propose a Relation-embedded Representation Reconstruction Network (R$^3$Net) to explicitly distinguish the real change from the large amount of clutter and irrelevant changes. Specifically, a relation-embedded module is first devised to explore potential changed objects in the large amount of clutter. Then, based on the semantic similarities of corresponding locations in the two images, a representation reconstruction module (RRM) is designed to learn the reconstruction representation and further model the difference representation. Besides, we introduce a syntactic skeleton predictor (SSP) to enhance the semantic interaction between change localization and caption generation. Extensive experiments show that the proposed method achieves the state-of-the-art results on two public datasets.

* Accepted by EMNLP 2021

View paper on

Share this with someone who'll enjoy it:

Title:R$^3$Net:Relation-embedded Representation Reconstruction Network for Change Captioning

Paper and Code