Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
Bing Cao, Quanhao Lu, Jiekang Feng, Qilong Wang, Pengfei Zhu, Qinghua Hu
Abstract
The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To empower the model's representation ability on density regression, we develop a new ensity-mbedded asked mdeling () method, which first takes the density map as an auxiliary modality to perform multimodal self-representation learning for image and density map. Although contributes to effective cross-modal regression guidance, it also brings in redundant background information, making it difficult to focus on the foreground regions. To handle this dilemma, we propose an efficient spatial adaptive masking derived from density maps to boost efficiency. Meanwhile, we employ an optical flow-based temporal collaborative fusion strategy to effectively capture the dynamic variations across frames, aligning features to derive multi-frame density residuals. The counting accuracy of the current frame is boosted by harnessing the information from adjacent frames. In addition, considering that most existing datasets are limited to human-centric scenarios, we first propose a large video bird counting dataset, DroneBird, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our DroneBird validate our superiority against the counterparts. The code and dataset are available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Related papers
- Video Individual Counting for Moving DronesYaowu Fan, Jia Wan, Tao Han, Antoni B. Chan et al.ICCV 2025 · 1 citation
- Data-Efficient Masked Video Modeling for Self-supervised Action RecognitionQiankun Li, Xiaolong Huang, Zhifan Wan, Lanqing Hu et al.ACM MM 2023 · 9 citations
- Detection, Tracking, and Counting Meets Drones in Crowds: A BenchmarkLongyin Wen, Dawei Du, Pengfei Zhu, Qinghua Hu et al.CVPR 2021
- Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingYuhang He, Zhiheng Ma, Xing Wei, Xiaopeng Hong et al.AAAI 2021 · 34 citations
- DenseTrack: Drone-Based Crowd Tracking via Density-Aware Motion-Appearance SynergyYi Lei, Huilin Zhu, Jingling Yuan, Guangli Xiang et al.ACM MM 2024 · 3 citations
