Decoupling What to Count and Where to See for Referring Expression Counting
Yuda Zou, Zijian Zhang, Yongchao Xu
摘要
Referring Expression Counting (REC) extends class-level object counting to the fine-grained subclass-level, aiming to enumerate objects matching a textual expression that specifies both the class and distinguishing attribute. A fundamental challenge, however, has been overlooked: annotation points are typically placed on class-representative locations (e.g., heads), forcing models to focus on class-level features while neglecting attribute information from other visual regions (e.g., legs for "walking"). To address this, we propose W2-Net, a novel framework that explicitly decouples the problem into "what to count" and "where to see" via a dual-query mechanism. Specifically, alongside the standard what-to-count (w2c) queries that localize the object, we introduce dedicated where-to-see (w2s) queries. The w2s queries are guided to seek and extract features from attributespecific visual regions, enabling precise subclass discrimination. Furthermore, we introduce Subclass Separable Matching (SSM), a novel matching strategy that incorporates a repulsive force to enhance inter-subclass separability during label assignment. W2-Net significantly outperforms the stateof-the-art on the REC-8K dataset, reducing counting error by 22.5% (validation) and 18.0% (test), and improving localization F1 by 7% and 8%, respectively. Code will be available.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper29
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Bayesian Loss for Crowd Count Estimation With Point SupervisionZhiheng Ma, Xing Wei, Xiaopeng Hong, Yihong GongICCV 2019 · 被引用 612 次
- Rethinking Counting and Localization in Crowds: A Purely Point-Based FrameworkQingyu Song, Changan Wang, Zhengkai Jiang, Yabiao Wang 等ICCV 2021 · 被引用 376 次
- Perspective-Guided Convolution Networks for Crowd CountingZhaoyi Yan, Yuchen Yuan, Wangmeng Zuo, Xiao Tan 等ICCV 2019 · 被引用 209 次
- CountGD: Multi-Modal Open-World CountingNiki Amini-Naieni, Tengda Han, Andrew ZissermanNeurIPS 2024 · 被引用 96 次
相关 Paper
- DCount: Decoupled Spatial Perception and Attribute Discrimination for Referring Expression CountingMing Li, Yupeng Hu, Yinwei Wei, Hao Liu 等ACM MM 2025 · 被引用 3 次
- Referring Expression CountingSiyang Dai, Jun Liu, Ngai-Man CheungCVPR 2024
- Revisiting Counterfactual Problems in Referring Expression ComprehensionZhihan Yu, Ruifan LiCVPR 2024 · 被引用 6 次
- Correspondence Matters for Video Referring Expression ComprehensionMeng Cao, Ji Jiang, Long Chen, Yuexian ZouACM MM 2022 · 被引用 10 次
- CoHD: A Counting-Aware Hierarchical Decoding Framework for Generalized Referring Expression SegmentationZhuoyan Luo, Yinghao Wu, Tianheng Cheng, Yong Liu 等ICCV 2025 · 被引用 1 次
