Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Apr 18, 2024

Han Fang, Xianghao Zang, Chao Ban, Zerun Feng, Lanxiang Zhou, Zhongjiang He, Yongxiang Li, Hao Sun

Figure 1 for ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Figure 2 for ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Figure 3 for ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Figure 4 for ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Share this with someone who'll enjoy it:

Abstract:Text-video retrieval aims to find the most relevant cross-modal samples for a given query. Recent methods focus on modeling the whole spatial-temporal relations. However, since video clips contain more diverse content than captions, the model aligning these asymmetric video-text pairs has a high risk of retrieving many false positive results. In this paper, we propose Probabilistic Token Aggregation (\textit{ProTA}) to handle cross-modal interaction with content asymmetry. Specifically, we propose dual partial-related aggregation to disentangle and re-aggregate token representations in both low-dimension and high-dimension spaces. We propose token-based probabilistic alignment to generate token-level probabilistic representation and maintain the feature representation diversity. In addition, an adaptive contrastive loss is proposed to learn compact cross-modal distribution space. Based on extensive experiments, \textit{ProTA} achieves significant improvements on MSR-VTT (50.9%), LSMDC (25.8%), and DiDeMo (47.2%).

View paper on

Share this with someone who'll enjoy it:

Title:ProTA: Probabilistic Token Aggregation for Text-Video Retrieval

Paper and Code