Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

Mar 09, 2024

Xiao Wang, Ju Huang, Shiao Wang, Chuanming Tang, Bo Jiang, Yonghong Tian, Jin Tang, Bin Luo

Figure 1 for Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

Figure 2 for Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

Figure 3 for Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

Figure 4 for Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

Share this with someone who'll enjoy it:

Abstract:Current event-/frame-event based trackers undergo evaluation on short-term tracking datasets, however, the tracking of real-world scenarios involves long-term tracking, and the performance of existing tracking algorithms in these scenarios remains unclear. In this paper, we first propose a new long-term and large-scale frame-event single object tracking dataset, termed FELT. It contains 742 videos and 1,594,474 RGB frames and event stream pairs and has become the largest frame-event tracking dataset to date. We re-train and evaluate 15 baseline trackers on our dataset for future works to compare. More importantly, we find that the RGB frames and event streams are naturally incomplete due to the influence of challenging factors and spatially sparse event flow. In response to this, we propose a novel associative memory Transformer network as a unified backbone by introducing modern Hopfield layers into multi-head self-attention blocks to fuse both RGB and event data. Extensive experiments on both FELT and RGB-T tracking dataset LasHeR fully validated the effectiveness of our model. The dataset and source code can be found at \url{https://github.com/Event-AHU/FELT_SOT_Benchmark}.

* In Peer Review

View paper on

Share this with someone who'll enjoy it:

Title:Long-term Frame-Event Visual Tracking: Benchmark Dataset and Baseline

Paper and Code