Prototype-Guided Dual-Transformer Reasoning for Video Individual Counting
Rui Li, Yishu Liu, Huafeng Li, Jinxing Li, Guangming Lu
Abstract
Video Individual Counting (VIC), which focuses on accurately tallying the total number of individuals in a video without duplication, is crucial for urban public space management and densely-populated areas planning. Existing methods suffer from limitations in terms of expensive manual annotation, and the efficiency of location or detection algorithms. In this work, we contribute a novel Prototype-guided Dual-Transformer Reasoning framework, termed PDTR, which takes both similarity and difference of adjacent frames into account to achieve accurate counting in an end-to-end regression manner. Specifically, we first design a multi-receptive field feature fusion module to acquire initial comprehensive representations. Subsequently, the dynamic prototype generation module memorizes consistent representations of similar information to generate prototypes. Additionally, to further dig out the shared and private features from different frames, a prototype cross-guided decoder and a privacy-decoupling module are designed. Extensive experiments conducted on two existing VIC datasets, consistently demonstrate the superiority of PDTR over state-of-the-art baselines.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 163b7e45-5db5-41e5-83d8-720b5b0b1cf6Cited by top-tier papers1
Ask how each one uses itRelated papers
- DR.VIC: Decomposition and Reasoning for Video Individual CountingTao Han, Lei Bai, Junyu Gao, Qi Wang et al.CVPR 2022 · 18 citations
- Weakly Supervised Video Individual CountingXinyan Liu, Guorong Li, Yuankai Qi, Ziheng Yan et al.CVPR 2024
- TargetVAU: Multimodal Anomaly-Aware Reasoning for Target Behavior Understanding in VideosLingru Zhou, Peng Wu, Manqing Zhang, Qingsheng Wang et al.AAAI 2026
- Convolutional Transformer based Dual Discriminator Generative Adversarial Networks for Video Anomaly DetectionXinyang Feng, Dongjin Song, Yuncong Chen, Zhengzhang Chen et al.ACM MM 2021 · 101 citations
- VideoTrack: Learning to Track Objects via Video TransformerFei Xie, Lei Chu, Jiahao Li, Yan Lu et al.CVPR 2023
