Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Apr 08, 2019

Peratham Wiriyathammabhum, Abhinav Shrivastava, Vlad I. Morariu, Larry S. Davis

Figure 1 for Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Figure 2 for Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Figure 3 for Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Figure 4 for Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Share this with someone who'll enjoy it:

Abstract:This paper presents a new task, the grounding of spatio-temporal identifying descriptions in videos. Previous work suggests potential bias in existing datasets and emphasizes the need for a new data creation schema to better model linguistic structure. We introduce a new data collection scheme based on grammatical constraints for surface realization to enable us to investigate the problem of grounding spatio-temporal identifying descriptions in videos. We then propose a two-stream modular attention network that learns and grounds spatio-temporal identifying descriptions based on appearance and motion. We show that motion modules help to ground motion-related words and also help to learn in appearance modules because modular neural networks resolve task interference between modules. Finally, we propose a future challenge and a need for a robust system arising from replacing ground truth visual annotations with automatic video object detector and temporal event localization.

View paper on

Share this with someone who'll enjoy it:

Title:Referring to Objects in Videos using Spatio-Temporal Identifying Descriptions

Paper and Code