Leveraging Debiased Cross-Modal Attention Maps and Code-Based Reasoning for Zero-Shot Referring Expression Comprehension
Juntao Chen, Wen Shen, Zhihua Wei, Lijun Sun, Hongyun Zhang
Abstract
Zero-shot Referring Expression Comprehension (REC) aims at locating an object described by a natural language query without training on task-specific datasets. Current approaches often utilize Vision-Language Models (VLMs) to perform region-text matching based on region proposals. However, this may downgrade their performance since VLMs often fail in relation understanding and isolated proposals inevitably lack global image context. To tackle these challenges, we first design a general formulation for code-based relation reasoning. It instructs Large Language Models (LLMs) to decompose complex relations and adaptively implement code for spatial and relation computation. Moreover, we directly extract region-text relevance from cross-modal attention maps in VLMs. Observing the inherent bias in VLMs, we further develop a simple yet effective bias deduction method, which enhances attention maps' capability to align text with the corresponding regions. Experimental results on four representative datasets demonstrate the SOTA performance of our method. On the RefCOCO dataset centered on spatial understanding, our method gets an average improvement of 10% over the previous zero-shot SOTA.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3743011c-1dfb-4f95-a1ce-384361203ea7Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Toolformer: Language Models Can Teach Themselves to Use ToolsTimo Schick, Jane Dwivedi-Yu, Roberto Dessì, Roberta Raileanu et al.NeurIPS 2023 · 5,989 citations
- Align before Fuse: Vision and Language Representation Learning with Momentum DistillationJunnan Li, Ramprasaath R. Selvaraju, Akhilesh Gotmare, Shafiq R. Joty et al.NeurIPS 2021 · 2,985 citations
Related papers
- Zero-Shot Referring Expression Comprehension via Structural Similarity Between Images and CaptionsZeyu Han, Fangrui Zhu, Qianru Lao, Huaizu JiangCVPR 2024
- From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression ComprehensionLihong Huang, Sheng-hua Zhong, Zhi Zhang, Yan LiuAAAI 2026
- VIRO: Robust and Efficient Neuro-Symbolic Reasoning with Verification for Referring Expression ComprehensionHyejin Park, Junhyuk Kwon, Suha Kwak, Jungseul OkCVPR 2026 · 1 citation
- ReCLIP: A Strong Zero-Shot Baseline for Referring Expression ComprehensionSanjay Subramanian, William Merrill, Trevor Darrell, Matt Gardner et al.ACL 2022 · 172 citations
- Zero-shot Visual Relation Detection via Composite Visual Cues from Large Language ModelsLin Li, Jun Xiao, Guikun Chen, Jian Shao et al.NeurIPS 2023 · 52 citations
