Efficient Masked AutoEncoder for Video Object Counting and A Large-Scale Benchmark
Bing Cao, Quanhao Lu, Jiekang Feng, Qilong Wang, Pengfei Zhu, Qinghua Hu
摘要
The dynamic imbalance of the fore-background is a major challenge in video object counting, which is usually caused by the sparsity of target objects. This remains understudied in existing works and often leads to severe under-/over-prediction errors. To tackle this issue in video object counting, we propose a density-embedded Efficient Masked Autoencoder Counting (E-MAC) framework in this paper. To empower the model's representation ability on density regression, we develop a new ensity-mbedded asked mdeling () method, which first takes the density map as an auxiliary modality to perform multimodal self-representation learning for image and density map. Although contributes to effective cross-modal regression guidance, it also brings in redundant background information, making it difficult to focus on the foreground regions. To handle this dilemma, we propose an efficient spatial adaptive masking derived from density maps to boost efficiency. Meanwhile, we employ an optical flow-based temporal collaborative fusion strategy to effectively capture the dynamic variations across frames, aligning features to derive multi-frame density residuals. The counting accuracy of the current frame is boosted by harnessing the information from adjacent frames. In addition, considering that most existing datasets are limited to human-centric scenarios, we first propose a large video bird counting dataset, DroneBird, in natural scenarios for migratory bird protection. Extensive experiments on three crowd datasets and our DroneBird validate our superiority against the counterparts. The code and dataset are available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Video Individual Counting for Moving DronesYaowu Fan, Jia Wan, Tao Han, Antoni B. Chan 等ICCV 2025 · 被引用 1 次
- Data-Efficient Masked Video Modeling for Self-supervised Action RecognitionQiankun Li, Xiaolong Huang, Zhifan Wan, Lanqing Hu 等ACM MM 2023 · 被引用 9 次
- Detection, Tracking, and Counting Meets Drones in Crowds: A BenchmarkLongyin Wen, Dawei Du, Pengfei Zhu, Qinghua Hu 等CVPR 2021
- Error-Aware Density Isomorphism Reconstruction for Unsupervised Cross-Domain Crowd CountingYuhang He, Zhiheng Ma, Xing Wei, Xiaopeng Hong 等AAAI 2021 · 被引用 34 次
- DenseTrack: Drone-Based Crowd Tracking via Density-Aware Motion-Appearance SynergyYi Lei, Huilin Zhu, Jingling Yuan, Guangli Xiang 等ACM MM 2024 · 被引用 3 次
