Point-Query Quadtree for Crowd Counting, Localization, and More
Chengxin Liu, Hao Lu, Zhiguo Cao, Tongliang Liu
Abstract
We show that crowd counting can be viewed as a decomposable point querying process. This formulation enables arbitrary points as input and jointly reasons whether the points are crowd and where they locate. The querying processing, however, raises an underlying problem on the number of necessary querying points. Too few imply underestimation; too many increase computational overhead. To address this dilemma, we introduce a decomposable structure, i.e., the point-query quadtree, and propose a new counting model, termed Point quEry Transformer (PET). PET implements decomposable point querying via data-dependent quadtree splitting, where each querying point could split into four new points when necessary, thus enabling dynamic processing of sparse and dense regions. Such a querying process yields an intuitive, universal modeling of crowd as both the input and output are interpretable and steerable. We demonstrate the applications of PET on a number of crowd-related tasks, including fully-supervised crowd counting and localization, partial annotation learning, and point annotation refinement, and also report state-of-the-art performance. For the first time, we show that a single counting model can address multiple crowd-related tasks across different learning paradigms. Code is available at https://github.com/cxliu0/PET.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b94e15cf-8da8-46ce-879c-573511dc4800Cited by top-tier papers14
- CrowdDiff: Multi-Hypothesis Crowd Density Estimation Using Diffusion ModelsYasiru Ranasinghe, Nithin Gopalakrishnan Nair, Wele Gedara Chaminda Bandara, Vishal M. PatelCVPR 2024 · 19 citations
- Enhancing Zero-Shot Object Counting via Text-Guided Local Ranking and Number-Evoked Global AttentionShiwei Zhang, Qi Zhou, Wei KeICCV 2025 · 7 citations
- Boosting Quantitive and Spatial Awareness for Zero-Shot Object CountingDa Zhang, Bingyu Li, Feiyu Wang, Zhiyuan Zhao et al.CVPR 2026 · 6 citations
- Embodied Crowd CountingRunling Long, Yunlong Wang, Jia Wan, Xiang Deng et al.NeurIPS 2025 · 3 citations
- Bootstrapping MLLM for Weakly‑Supervised Class‑Agnostic Object CountingXiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao et al.ICLR 2026 · 3 citations
Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 612 citations
- Distribution Matching for Crowd CountingBoyu Wang, Huidong Liu, Dimitris Samaras, Minh Hoai NguyenNeurIPS 2020 · 443 citations
- Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkQingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang et al.ICCV 2021 · 376 citations
Related papers
- End-to-End Multi-Person Pose Estimation with TransformersDahu Shi, Xing Wei, Liangqi Li, Ye Ren et al.CVPR 2022 · 147 citations
- Semi-supervised Crowd Counting via Density AgencyHui Lin, Zhiheng Ma, Xiaopeng Hong, Yaowei Wang et al.ACM MM 2022 · 37 citations
- Stratified Transformer for 3D Point Cloud SegmentationXin Lai, Jianhui Liu, Li Jiang, Liwei Wang et al.CVPR 2022 · 494 citations
- A Hierarchical Spatial Transformer for Massive Point Samples in Continuous SpaceWenchong He, Zhe Jiang, Tingsong Xiao, Zelin Xu et al.NeurIPS 2023 · 20 citations
- Group Pose: A Simple Baseline for End-to-End Multi-person Pose EstimationHuan Liu, Qiang Chen, Zichang Tan, Jiang-Jiang Liu et al.ICCV 2023 · 50 citations
