Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Bootstrapping Referring Multi-Object Tracking

Jun 07, 2024

Yani Zhang, Dongming Wu, Wencheng Han, Xingping Dong

Figure 1 for Bootstrapping Referring Multi-Object Tracking

Figure 2 for Bootstrapping Referring Multi-Object Tracking

Figure 3 for Bootstrapping Referring Multi-Object Tracking

Figure 4 for Bootstrapping Referring Multi-Object Tracking

Share this with someone who'll enjoy it:

Abstract:Referring multi-object tracking (RMOT) aims at detecting and tracking multiple objects following human instruction represented by a natural language expression. Existing RMOT benchmarks are usually formulated through manual annotations, integrated with static regulations. This approach results in a dearth of notable diversity and a constrained scope of implementation. In this work, our key idea is to bootstrap the task of referring multi-object tracking by introducing discriminative language words as much as possible. In specific, we first develop Refer-KITTI into a large-scale dataset, named Refer-KITTI-V2. It starts with 2,719 manual annotations, addressing the issue of class imbalance and introducing more keywords to make it closer to real-world scenarios compared to Refer-KITTI. They are further expanded to a total of 9,758 annotations by prompting large language models, which create 617 different words, surpassing previous RMOT benchmarks. In addition, the end-to-end framework in RMOT is also bootstrapped by a simple yet elegant temporal advancement strategy, which achieves better performance than previous approaches. The source code and dataset is available at https://github.com/zyn213/TempRMOT.

View paper on

Share this with someone who'll enjoy it:

Title:Bootstrapping Referring Multi-Object Tracking

Paper and Code