Get our free extension to see links to code for papers anywhere online!Free add-on: code for papers everywhere!Free add-on: See code for papers anywhere!

Add to Chrome

Add to Firefox

Add to Edge

Title:Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Nov 20, 2024

Bing Cao, Quanhao Lu, Jiekang Feng, Pengfei Zhu, Qinghua Hu, Qilong Wang

Figure 1 for Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Figure 2 for Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Figure 3 for Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Figure 4 for Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Share this with someone who'll enjoy it:

Abstract:The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of foreground objects. This often leads to severe under- and over-prediction problems and has been less studied in existing works. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To effectively capture the dynamic variations across frames, we utilize an optical flow-based temporal collaborative fusion that aligns features to derive multi-frame density residuals. The counting accuracy of the current frame is boosted by harnessing the information from adjacent frames. More importantly, to empower the representation ability of dynamic foreground objects for intra-frame, we first take the density map as an auxiliary modality to perform $\mathtt{D}$ensity-$\mathtt{E}$mbedded $\mathtt{M}$asked m$\mathtt{O}$deling ($\mathtt{DEMO}$) for multimodal self-representation learning to regress density map. However, as $\mathtt{DEMO}$ contributes effective cross-modal regression guidance, it also brings in redundant background information and hard to focus on foreground regions. To handle this dilemma, we further propose an efficient spatial adaptive masking derived from density maps to boost efficiency. In addition, considering most existing datasets are limited to human-centric scenarios, we first propose a large video bird counting dataset $\textit{DroneBird}$, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our $\textit{DroneBird}$ validate our superiority against the counterparts.

View paper on

Share this with someone who'll enjoy it:

Title:Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark

Paper and Code