Decoupling What to Count and Where to See for Referring Expression Counting
Yuda Zou, Zijian Zhang, Yongchao Xu
Abstract
Referring Expression Counting (REC) extends class-level object counting to the fine-grained subclass-level, aiming to enumerate objects matching a textual expression that specifies both the class and distinguishing attribute. A fundamental challenge, however, has been overlooked: annotation points are typically placed on class-representative locations (e.g., heads), forcing models to focus on class-level features while neglecting attribute information from other visual regions (e.g., legs for "walking"). To address this, we propose W2-Net, a novel framework that explicitly decouples the problem into "what to count" and "where to see" via a dual-query mechanism. Specifically, alongside the standard what-to-count (w2c) queries that localize the object, we introduce dedicated where-to-see (w2s) queries. The w2s queries are guided to seek and extract features from attributespecific visual regions, enabling precise subclass discrimination. Furthermore, we introduce Subclass Separable Matching (SSM), a novel matching strategy that incorporates a repulsive force to enhance inter-subclass separability during label assignment. W2-Net significantly outperforms the stateof-the-art on the REC-8K dataset, reducing counting error by 22.5% (validation) and 18.0% (test), and improving localization F1 by 7% and 8%, respectively. Code will be available.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 305c0933-1c7c-4adf-9f61-c624c3df85e2Builds on29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 612 citations
- Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkQingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang et al.ICCV 2021 · 376 citations
- Perspective-Guided Convolution Networks for Crowd CountingZhaoyi Yan, Yuchen Yuan, Wangmeng Zuo, Xiao Tan et al.ICCV 2019 · 209 citations
- CountGD: Multi-Modal Open-World CountingNiki Amini-Naieni, Tengda Han, Andrew ZissermanNeurIPS 2024 · 96 citations
Related papers
- DCount: Decoupled Spatial Perception and Attribute Discrimination for Referring Expression CountingMing Li, Yupeng Hu, Yinwei Wei, Hao Liu et al.ACM MM 2025 · 3 citations
- Referring Expression CountingSiyang Dai, Jun Liu, Ngai-Man CheungCVPR 2024
- Revisiting Counterfactual Problems in Referring Expression ComprehensionZhihan Yu, Ruifan LiCVPR 2024 · 6 citations
- Correspondence Matters for Video Referring Expression ComprehensionMeng Cao, Ji Jiang, Long Chen, Yuexian ZouACM MM 2022 · 10 citations
- CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression SegmentationZhuoyan Luo, Yinghao Wu, Tianheng Cheng, Yong Liu et al.ICCV 2025 · 1 citation
