Weakly Supervised Video Individual Counting
Xinyan Liu, Guorong Li, Yuankai Qi, Ziheng Yan, Zhenjun Han, Anton van den Hengel, Ming-Hsuan Yang, Qingming Huang
Abstract
Video Individual Counting (VIC) aims to predict the number of unique individuals in a single video. Existing methods learn representations based on trajectory labels for individuals, which are annotation-expensive. To provide a more realistic reflection of the underlying practical challenge, we introduce a weakly supervised VIC task, wherein trajectory labels are not provided. Instead, two types of labels are provided to indicate traffic entering the field of view (inflow) and leaving the field view (outflow). We also propose the first solution as a baseline that formulates the task as a weakly supervised contrastive learning problem under group-level matching. In doing so, we devise an end-to-end trainable soft contrastive loss to drive the network to distinguish inflow, outflow, and the remaining. To facilitate future study in this direction, we generate annotations from the existing VIC datasets SenseCrowd and CroHD and also build a new dataset, UAVVIC. Extensive results show that our baseline weakly supervised method outperforms supervised methods, and thus, little information is lost in the transition to the more practically relevant weakly supervised task. The code and trained model can be found at CGNet.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2d3703e-5ad5-4307-a5b8-629912905ff8Cited by top-tier papers1
Ask how each one uses itBuilds on6
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- SMILEtrack: SiMIlarity LEarning for Occlusion-Aware Multiple Object TrackingYu-Hsiang Wang, Jun-Wei Hsieh, Ping-Yang Chen, Ming-Ching Chang et al.AAAI 2024 · 96 citations
- CWCL: Cross-Modal Transfer with Continuously Weighted Contrastive LossRakshith Sharma Srinivasa, Jaejin Cho, Chouchang Yang, Yashas Malur Saidutta et al.NeurIPS 2023 · 25 citations
- Discriminative Appearance Modeling With Multi-Track Pooling for Real-Time Multi-Object TrackingChanho Kim, Fuxin Li, Mazen Alotaibi, James M. RehgCVPR 2021
- Tracking Pedestrian Heads in Dense CrowdRamana Sundararaman, Cedric De Almeida Braga, Éric Marchand, Julien PettréCVPR 2021
Related papers
- Flowing Crowd to Count Flows: A Self-Supervised Framework for Video Individual CountingFeng-Kai Huang, Bo-Lun Huang, Li-Wu Tsao, Jhih-Ciang Wu et al.ACM MM 2025 · 1 citation
- Video Individual Counting for Moving DronesYaowu Fan, Jia Wan, Tao Han, Antoni B. Chan et al.ICCV 2025 · 1 citation
- DR.VIC: Decomposition and Reasoning for Video Individual CountingTao Han, Lei Bai, Junyu Gao, Qi Wang et al.CVPR 2022 · 18 citations
- Prototype-Guided Dual-Transformer Reasoning for Video Individual CountingRui Li, Yishu Liu, Huafeng Li, Jinxing Li et al.ACM MM 2024 · 2 citations
- TLMA: Mitigating the Impact of Weakly Labeled Information for Video Anomaly DetectionRong Xu, Runqi Wang, Yingjun Zhang, Tao Tao et al.CVPR 2026
