From Pixels to Logic: A Perception-Reasoning Decomposition Framework for Open-World Referring Expression Comprehension
Lihong Huang, Sheng-hua Zhong, Zhi Zhang, Yan Liu
Abstract
Recent advances in Referring Expression Comprehension (REC) have been largely driven by supervised learning on curated datasets, where each expression is assumed to refer to exactly one known object. However, such assumptions rarely hold in real-world scenarios, where expressions can refer to multiple objects, fail to refer to any, or involve novel categories and complex semantics. These challenges define the task of open-world REC, which demands robust generalization and structured reasoning beyond the scope of traditional REC methods. In this work, we introduce a novel, trainingfree framework that decouples visual perception from linguistic reasoning to address open-world REC. Our method first transforms the visual scene into a rich textual representation using an open-vocabulary multimodal perception module. It then employs a reasoning language model to interpret the referring expression and perform explicit logical inference over the perceived scene, enabling transparent decision-making and strong generalization in open-world scenarios. Experiments on three standard REC benchmarks as well as two more challenging ones, gRefCOCO and D 3 , demonstrate that our framework achieves highly competitive zero-shot performance, often surpassing supervised baselines.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0336ee8d-0d4c-4d07-94be-513de78904d3Builds on27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Towards Universal Perception through Language-Guided Open-World Object DetectionZihan Wang, Yunhang Shen, Yuan Fang, Zuwei Long et al.ACM MM 2025 · 1 citation
- Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionZhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong et al.CVPR 2020
- Leveraging Debiased Cross-Modal Attention Maps and Code-Based Reasoning for Zero-Shot Referring Expression ComprehensionJuntao Chen, Wen Shen, Zhihua Wei, Lijun Sun et al.ICCV 2025 · 1 citation
- Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression TasksQihua Dong, Kuo Yang, Lin Ju, Handong Zhao et al.ICLR 2026 · 13 citations
- Towards Further Comprehension on Referring Expression with RationaleRengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang et al.ACM MM 2022 · 2 citations
