Learning Cross-Modal Context Graph for Visual Grounding
Yongfei Liu, Bo Wan, Xiaodan Zhu, Xuming He
摘要
Visual grounding is a ubiquitous building block in many vision-language tasks and yet remains challenging due to large variations in visual and linguistic features of grounding entities, strong context effect and the resulting semantic ambiguities. Prior works typically focus on learning representations of individual phrases with limited context information. To address their limitations, this paper proposes a language-guided graph representation to capture the global context of grounding entities and their relations, and develop a cross-modal graph matching strategy for the multiple-phrase visual grounding task. In particular, we introduce a modular graph neural network to compute context-aware representations of phrases and object proposals respectively via message propagation, followed by a graph-based matching module to generate globally consistent localization of grounding phrases. We train the entire graph neural network jointly in a two-stage strategy and evaluate it on the Flickr30K Entities benchmark. Extensive experiments show that our method outperforms the prior state of the arts by a sizable margin, evidencing the efficacy of our grounding framework. Code is available at https://github.com/youngfly11/LCMCG-PyTorch.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- What does CLIP know about a red circle? Visual prompt engineering for VLMsAleksandar Shtedritski, Christian Rupprecht, Andrea VedaldiICCV 2023 · 被引用 262 次
- Free-form Description Guided 3D Visual Graph Network for Object Grounding in Point CloudMingtao Feng, Zhen Li, Qi Li, Liang Zhang 等ICCV 2021 · 被引用 115 次
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang 等CVPR 2022 · 被引用 89 次
- Look Around and Refer: 2D Synthetic Semantics Knowledge Distillation for 3D Visual GroundingEslam Mohamed Bakr, Yasmeen Alsaedy, Mohamed ElhoseinyNeurIPS 2022 · 被引用 66 次
- Disentangled Motif-aware Graph Learning for Phrase GroundingZongshen Mu, Siliang Tang, Jie Tan, Qiang Yu 等AAAI 2021 · 被引用 38 次
它引用的顶会 Paper1
相关 Paper
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu 等ACM MM 2020 · 被引用 35 次
- G3raphGround: Graph-Based Language GroundingMohit Bajaj, Lanjun Wang, Leonid SigalICCV 2019 · 被引用 67 次
- Relation-aware Instance Refinement for Weakly Supervised Visual GroundingYongfei Liu, Bo Wan, Lin Ma, Xuming HeCVPR 2021
- Improving Visual Grounding with Visual-Linguistic Verification and Iterative ReasoningLi Yang, Yan Xu, Chunfeng Yuan, Wei Liu 等CVPR 2022 · 被引用 146 次
- Improving Zero-Shot Phrase Grounding via Reasoning on External Knowledge and Spatial RelationsZhan Shi, Yilin Shen, Hongxia Jin, Xiaodan ZhuAAAI 2022 · 被引用 7 次
