Graph-Structured Referring Expression Reasoning in the Wild
Sibei Yang, Guanbin Li, Yizhou Yu
Abstract
Grounding referring expressions aims to locate in an image an object referred to by a natural language expression. The linguistic structure of a referring expression provides a layout of reasoning over the visual contents, and it is often crucial to align and jointly understand the image and the referring expression. In this paper, we propose a scene graph guided modular network (SGMN), which performs reasoning over a semantic graph and a scene graph with neural modules under the guidance of the linguistic structure of the expression. In particular, we model the image as a structured semantic graph, and parse the expression into a language scene graph. The language scene graph not only decodes the linguistic structure of the expression, but also has a consistent representation with the image semantic graph. In addition to exploring structured solutions to grounding referring expressions, we also propose Ref-Reasoning, a large-scale real-world dataset for structured referring expression reasoning. We automatically generate referring expressions over the scene graphs of images using diverse expression templates and functional programs. This dataset is equipped with real-world visual contents as well as semantically rich expressions with different reasoning layouts. Experimental results show that our SGMN 1 not only significantly outperforms existing state-of-the-art algorithms on the new Ref-Reasoning dataset, but also surpasses state-of-the-art structured methods on commonly used benchmark datasets. It can also provide interpretable visual evidences of reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 220f3e68-8885-4cb7-a945-cd0e280e3c9aCited by top-tier papers31
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou et al.ICCV 2021 · 468 citations
- Text-Guided Graph Neural Networks for Referring 3D Instance SegmentationPin-Hao Huang, Han-Hung Lee, Hwann-Tzong Chen, Tyng-Luh LiuAAAI 2021 · 191 citations
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu et al.CVPR 2022 · 146 citations
- Ref-NMS: Breaking Proposal Bottlenecks in Two-Stage Referring Expression GroundingLong Chen, Wenbo Ma, Jun Xiao, Hanwang Zhang et al.AAAI 2021 · 118 citations
- Debiased Visual Question Answering from Feature and Sample PerspectivesZhiquan Wen, Guanghui Xu, Mingkui Tan, Qingyao Wu et al.NeurIPS 2021 · 102 citations
Builds on2
Related papers
- Cops-Ref: A New Dataset and Task on Compositional Referring Expression ComprehensionZhenfang Chen, Peng Wang, Lin Ma, Kwan-Yee K. Wong et al.CVPR 2020
- Look Around Before Locating: Considering Content and Structure Information for Visual GroundingShiyi Zheng, Peizhi Zhao, Zhilong Zheng, Peihang He et al.AAAI 2025 · 3 citations
- Exploring Logical Reasoning for Referring Expression ComprehensionYing Cheng, Ruize Wang, Jiashuo Yu, Rui-Wei Zhao et al.ACM MM 2021 · 12 citations
- Advancing Visual Grounding with Scene Knowledge: Benchmark and MethodZhihong Chen, Ruifei Zhang, Yibing Song, Xiang Wan et al.CVPR 2023
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu et al.ACM MM 2020 · 35 citations
