Regressor-Segmenter Mutual Prompt Learning for Crowd Counting
Mingyue Guo, Li Yuan, Zhaoyi Yan, Binghui Chen, Yaowei Wang, Qixiang Ye
Abstract
Crowd counting has achieved significant progress by training regressors to predict instance positions. In heavily crowded scenarios, however, regressors are challenged by uncontrollable annotation variance, which causes density map bias and context information inaccuracy. In this study, we propose mutual prompt learning (mPrompt), which leverages a regressor and a segmenter as guidance for each other, solving bias and inaccuracy caused by annotation variance while distinguishing foreground from background. In specific, mPrompt leverages point annotations to tune the segmenter and predict pseudo head masks in a way of point prompt learning. It then uses the predicted segmentation masks, which serve as spatial constraint, to rectify biased point annotations as context prompt learning. mPrompt defines a way of mutual information maximization from prompt learning, mitigating the impact of annotation variance while improving model accuracy. Experiments show that mPrompt significantly reduces the Mean Average Error (MAE), demonstrating the potential to be general framework for down-stream vision tasks. Code is available at https://github.com/csguomy/mPrompt .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- YOLO-Count: Differentiable Object Counting for Text-to-Image GenerationGuanning Zeng, Xiang Zhang, Zirui Wang, Haiyang Xu et al.ICCV 2025 · 4 citations
- Bootstrapping MLLM for Weakly‑Supervised Class‑Agnostic Object CountingXiaowen Zhang, Zijie Yue, Yong Luo, Cairong Zhao et al.ICLR 2026 · 3 citations
- Video Individual Counting for Moving DronesYaowu Fan, Jia Wan, Tao Han, Antoni B. Chan et al.ICCV 2025 · 1 citation
- CountSE: Soft Exemplar Open-Set Object CountingShuai Liu, Peng Zhang, Shiwei Zhang, Wei KeICCV 2025 · 1 citation
- Adapting Lightweight Image-based Counting Models for Video Crowd CountingWeibo Shu, Antoni B. ChanCVPR 2026
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order SensitivityYao Lu, Max Bartolo, Alastair Moore, Sebastian Riedel et al.ACL 2022 · 1,494 citations
- Towards a Unified View of Parameter-Efficient Transfer LearningJunxian He, Chunting Zhou, Xuezhe Ma, Taylor Berg-Kirkpatrick et al.ICLR 2022 · 1,182 citations
Related papers
- Semi-supervised Crowd Counting via Density AgencyHui Lin, Zhiheng Ma, Xiaopeng Hong, Yaowei Wang et al.ACM MM 2022 · 37 citations
- Proximal Mapping Loss: Understanding Loss Functions in Crowd Counting & LocalizationWei Lin, Jia Wan, Antoni B. ChanICLR 2025
- Single Domain Generalization for Crowd CountingZhuoxuan Peng, S.-H. Gary ChanCVPR 2024 · 27 citations
- A Fixed-Point Approach to Unified Prompt-Based CountingWei Lin, Antoni B. ChanAAAI 2024 · 11 citations
- PromptMoE: A Segmentation Refinement Framework Leveraging Mixture of Experts for Improved PromptingStephen Price, Danielle L. Cote, Elke A. RundensteinerCVPR 2026
