Whether you can locate or not? Interactive Referring Expression Generation
Fulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie Wang
Abstract
Referring Expression Generation (REG) aims to generate unambiguous Referring Expressions (REs) for objects in a visual scene, with a dual task of Referring Expression Comprehension (REC) to locate the referred object. Existing methods construct REG models independently by using only the REs as ground truth for model training, without considering the potential interaction between REG and REC models. In this paper, we propose an Interactive REG (IREG) model that can interact with a real REC model, utilizing signals indicating whether the object is located and the visual region located by the REC model to gradually modify REs. Our experimental results on three RE benchmark datasets, RefCOCO, RefCOCO+, and RefCOCOg show that IREG outperforms previous state-of-the-art methods on popular evaluation metrics. Furthermore, a human evaluation shows that IREG generates better REs with the capability of interaction 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 904120c4-e1e1-4712-b57b-b562214a8befCited by top-tier papers3
- ISR: Self-Refining Referring Expressions for Entity GroundingZhuocheng Yu, Bingchan Zhao, Yifan Song, Sujian Li et al.ACL 2025 · 1 citation
- Multi-Modal Instruction Tuned LLMs with Fine-Grained Visual PerceptionJunwen He, Yifan Wang, Lijun Wang, Huchuan Lu et al.CVPR 2024
- Breaking the Regional Perception Bottleneck of Multimodal Large Language Models via External Reasoning FrameworkJinrong Zhang, Zhaoyang Xu, Xusheng He, Xinrui Li et al.CVPR 2026
Builds on12
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- OFA: Unifying Architectures, Tasks, and Modalities Through a Simple Sequence-to-Sequence Learning FrameworkPeng Wang, An Yang, Rui Men, Junyang Lin et al.ICML 2022 · 1,058 citations
- Unifying Vision-and-Language Tasks via Text GenerationJaemin Cho, Jie Lei, Hao Tan, Mohit BansalICML 2021 · 624 citations
- A Fast and Accurate One-Stage Approach to Visual GroundingZhengyuan Yang, Boqing Gong, Liwei Wang, Wenbing Huang et al.ICCV 2019 · 441 citations
Related papers
- Towards Unifying Reference Expression Generation and ComprehensionDuo Zheng, Tao Kong, Ya Jing, Jiaan Wang et al.EMNLP 2022 · 4 citations
- Towards Further Comprehension on Referring Expression with RationaleRengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang et al.ACM MM 2022 · 2 citations
- Generation and Comprehension Hand-in-Hand: Vision-guided Expression Diffusion for Boosting Referring Expression Generation and ComprehensionJingcheng Ke, Jun-Cheng Chen, I-Hong Jhuo, Chia-Wen Lin et al.ICLR 2025
- Referring Expression Instance Retrieval and A Strong End-to-End BaselineXiangzhao Hao, Kuan Zhu, Hongyu Guo, Haiyun Guo et al.ACM MM 2025
- A Real-Time Cross-Modality Correlation Filtering Method for Referring Expression ComprehensionYue Liao, Si Liu, Guanbin Li, Fei Wang et al.CVPR 2020
