DCount: Decoupled Spatial Perception and Attribute Discrimination for Referring Expression Counting
Ming Li, Yupeng Hu, Yinwei Wei, Hao Liu, Haocong Wang, Weili Guan
Abstract
Referring Expression Counting (REC) is an emerging task that aims to count specific objects in images based on textual phrases describing their attributes and categories. While current REC baselines inherit architectures from pre-trained open-vocabulary object detectors and demonstrate promising counting and localization capabilities, they overlook critical limitations in the original single-decoder design with shared object queries. This architectural constraint entangles the semantic and localization perception processes, hindering fine-grained understanding of attribute-aware visual features. To address these challenges, we propose DCount, a decoupled counting framework comprising two innovative components: a Decoupled Dual-Decoder (DDD) module and an Attribute Semantic Discriminator (ASD) module. The DDD module separates spatial perception tasks by employing distinct semantic and localization decoders with task-specific object queries, thereby enhancing the capture of discriminative visual features. Building upon the positional and semantic feedback from DDD, the ASD module introduces a two-stage filtering strategy to explicitly mine challenging hard negative attribute samples in the visual domain, while synergistically refining attribute discrimination across both modalities through contrastive learning in the textual domain. Our method achieves state-of-the-art results on both the REC and Zero-Shot Object Counting (ZSOC) benchmarks.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get 09b936d9-9636-4de6-b647-b9599d18e3fcCited by top-tier papers1
Ask how each one uses itRelated papers
- Decoupling What to Count and Where to See for Referring Expression CountingYuda Zou, Zijian Zhang, Yongchao XuAAAI 2026
- Revisiting Counterfactual Problems in Referring Expression ComprehensionZhihan Yu, Ruifan LiCVPR 2024 · 6 citations
- Referring Expression CountingSiyang Dai, Jun Liu, Ngai-Man CheungCVPR 2024
- Latent Expression Generation for Referring Image Segmentation and GroundingSeonghoon Yu, Joonbeom Hong, Joonseok Lee, Jeany SonICCV 2025 · 1 citation
- From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression ComprehensionLihong Huang, Sheng-hua Zhong, Zhi Zhang, Yan LiuAAAI 2026
