Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Event-aware Video Corpus Moment Retrieval

Feb 21, 2024

Danyang Hou, Liang Pang, Huawei Shen, Xueqi Cheng

Figure 1 for Event-aware Video Corpus Moment Retrieval

Figure 2 for Event-aware Video Corpus Moment Retrieval

Figure 3 for Event-aware Video Corpus Moment Retrieval

Figure 4 for Event-aware Video Corpus Moment Retrieval

Share this with someone who'll enjoy it:

Abstract:Video Corpus Moment Retrieval (VCMR) is a practical video retrieval task focused on identifying a specific moment within a vast corpus of untrimmed videos using the natural language query. Existing methods for VCMR typically rely on frame-aware video retrieval, calculating similarities between the query and video frames to rank videos based on maximum frame similarity.However, this approach overlooks the semantic structure embedded within the information between frames, namely, the event, a crucial element for human comprehension of videos. Motivated by this, we propose EventFormer, a model that explicitly utilizes events within videos as fundamental units for video retrieval. The model extracts event representations through event reasoning and hierarchical event encoding. The event reasoning module groups consecutive and visually similar frame representations into events, while the hierarchical event encoding encodes information at both the frame and event levels. We also introduce anchor multi-head self-attenion to encourage Transformer to capture the relevance of adjacent content in the video. The training of EventFormer is conducted by two-branch contrastive learning and dual optimization for two sub-tasks of VCMR. Extensive experiments on TVR, ANetCaps, and DiDeMo benchmarks show the effectiveness and efficiency of EventFormer in VCMR, achieving new state-of-the-art results. Additionally, the effectiveness of EventFormer is also validated on partially relevant video retrieval task.

* 11 pages, 5 figures, 9 tables

View paper on

Share this with someone who'll enjoy it:

Title:Event-aware Video Corpus Moment Retrieval

Paper and Code