Revisiting Counterfactual Problems in Referring Expression Comprehension
Zhihan Yu, Ruifan Li
Abstract
Traditional referring expression comprehension (REC) aims to locate the target referent in an image guided by a text query. Several previous methods have studied on the Counterfactual problem in REC (C-REC) where the objects for a given query cannot be found in the image. However, these methods focus on the overall image-text or specific attribute mismatch only. In this paper, we address the C-REC problem from a deep perspective of fine- grained attributes. To this aim, we first propose a fine-grained counterfactual sample generation method to construct C-REC datasets. Specifically, we leverage pre-trained language model such as BERT to modify the attribute words in the queries, obtaining the corresponding counterfactual samples. Furthermore, we propose a C-REC framework. We first adopt three encoders to extract image, text and attribute features. Then, our dual-branch attentive fusion module fuses these cross-modal features with two branches by an attention mechanism. At last, two prediction heads generate a bounding box and a counterfactual label, respectively. In addition, we incorporate contrastive learning with the generated counterfactual samples as negatives to enhance the counterfactual perception. Extensive experiments show that our framework achieves promising performance on both public REC datasets RefCOCO/+lg and our constructed C-REC datasets C-RefCOCO/+lg. The code and data are available at https://github.com/Glacier0012/CREC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 25b65f02-d5e7-4de3-963d-14f167e73cdaCited by top-tier papers4
- Implicit Counterfactual Learning for Audio-Visual SegmentationMingfeng Zha, Tianyu Li, Guoyin Wang, Peng Wang et al.ICCV 2025 · 3 citations
- Referring to Any PersonQing Jiang, Lin Wu, Zhaoyang Zeng, Tianhe Ren et al.ICCV 2025 · 1 citation
- Diffusion-Assisted Progressive Learning for Weakly Supervised Phrase LocalizationPengyue Lin, Yanyang Hu, Xinjing Liu, Wenqi Jia et al.AAAI 2026
- Referring Expression Comprehension for Small ObjectsKanoko Goto, Takumi Hirose, Mahiro Ukai, Shuhei Kurita et al.ICCV 2025
Builds on14
- MDETR - Modulated Detection for End-to-End Multi-Modal UnderstandingAishwarya Kamath, Mannat Singh, Yann LeCun, Gabriel Synnaeve et al.ICCV 2021 · 1,114 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
- Learning to Assemble Neural Module Tree Networks for Visual GroundingDaqing Liu, Hanwang Zhang, Feng Wu, Zheng-Jun ZhaICCV 2019 · 317 citations
Related papers
- Referring Expression Instance Retrieval and A Strong End-to-End BaselineXiangzhao Hao, Kuan Zhu, Hongyu Guo, Haiyun Guo et al.ACM MM 2025
- Bottom-Up and Bidirectional Alignment for Referring Expression ComprehensionLiuwu Li, Yuqi Bu, Yi CaiACM MM 2021 · 11 citations
- RefCrowd: Grounding the Target in Crowd with Referring ExpressionsHeqian Qiu, Hongliang Li, Taijin Zhao, Lanxiao Wang et al.ACM MM 2022 · 10 citations
- DCount: Decoupled Spatial Perception and Attribute Discrimination for Referring Expression CountingMing Li, Yupeng Hu, Yinwei Wei, Hao Liu et al.ACM MM 2025 · 3 citations
- FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression ComprehensionJunzhuo Liu, Xuzheng Yang, Weiwei Li, Peng WangEMNLP 2024 · 2 citations
