Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

May 22, 2024

Yuting Wang, Jinpeng Wang, Bin Chen, Tao Dai, Ruisheng Luo, Shu-Tao Xia

Figure 1 for GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

Figure 2 for GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

Figure 3 for GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

Figure 4 for GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

Share this with someone who'll enjoy it:

Abstract:Given a text query, partially relevant video retrieval (PRVR) aims to retrieve untrimmed videos containing relevant moments. Due to the lack of moment annotations, the uncertainty lying in clip modeling and text-clip correspondence leads to major challenges. Despite the great progress, existing solutions either sacrifice efficiency or efficacy to capture varying and uncertain video moments. What's worse, few methods have paid attention to the text-clip matching pattern under such uncertainty, exposing the risk of semantic collapse. To address these issues, we present GMMFormer v2, an uncertainty-aware framework for PRVR. For clip modeling, we improve a strong baseline GMMFormer with a novel temporal consolidation module upon multi-scale contextual features, which maintains efficiency and improves the perception for varying moments. To achieve uncertainty-aware text-clip matching, we upgrade the query diverse loss in GMMFormer to facilitate fine-grained uniformity and propose a novel optimal matching loss for fine-grained text-clip alignment. Their collaboration alleviates the semantic collapse phenomenon and neatly promotes accurate correspondence between texts and moments. We conduct extensive experiments and ablation studies on three PRVR benchmarks, demonstrating remarkable improvement of GMMFormer v2 compared to the past SOTA competitor and the versatility of uncertainty-aware text-clip matching for PRVR. Code is available at \url{https://github.com/huangmozhi9527/GMMFormer_v2}.

View paper on

Share this with someone who'll enjoy it:

Title:GMMFormer v2: An Uncertainty-aware Framework for Partially Relevant Video Retrieval

Paper and Code