Towards Further Comprehension on Referring Expression with Rationale
Rengang Li, Baoyu Fan, Xiaochuan Li, Runze Zhang, Zhenhua Guo, Kun Zhao, Yaqian Zhao, Weifeng Gong, Endong Wang
Abstract
Referring Expression Comprehension (REC) is one important research branch in visual grounding, where the goal of REC is to localize a relevant object in the image, given an expression in the form of text to exactly describe a specific object. However, existing REC tasks aim at text content filtering and image object locating, which are evaluated based on the precision of the detection boxes. This may lead models to skip the learning process of multimodal comprehension directly and achieve good performance. In this paper, we work on how to enable an artificial agent to understand RE further and propose a more comprehensive task, called Further Comprehension on Referring Expression (FREC). In this task, we mainly focus on three sub-tasks: 1) correcting the erroneous text expression based on visual information; 2) generating the rationale of this input expression; 3) localizing the proper object based on the corrected expression. Accordingly, we make a new dataset named Further-RefCOCOs based on the RefCOCO, RefCOCO+, RefCOCOg benchmark datasets for this new task and make it publicly available. After that, we design a novel end-to-end pipeline to achieve these sub-tasks simultaneously. The experimental results demonstrate the validity of the proposed pipeline. We believe this work will motivate more researchers to explore along with this direction, and promote the development of visual grounding.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get bc7a984f-d980-411c-bb3d-e32cb14b5712Related papers
- FineCops-Ref: A new Dataset and Task for Fine-Grained Compositional Referring Expression ComprehensionJunzhuo Liu, Xuzheng Yang, Weiwei Li, Peng WangEMNLP 2024 · 2 citations
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 270 citations
- Ref-Adv: Exploring MLLM Visual Reasoning in Referring Expression TasksQihua Dong, Kuo Yang, Lin Ju, Handong Zhao et al.ICLR 2026 · 13 citations
- Whether you can locate or not? Interactive Referring Expression GenerationFulong Ye, Yuxing Long, Fangxiang Feng, Xiaojie WangACM MM 2023 · 6 citations
- Unveiling Parts Beyond Objects: Towards Finer-Granularity Referring Expression SegmentationWenxuan Wang, Tongtian Yue, Yisi Zhang, Longteng Guo et al.CVPR 2024 · 5 citations
