G3raphGround: Graph-Based Language Grounding
Mohit Bajaj, Lanjun Wang, Leonid Sigal
摘要
In this paper we present an end-to-end framework for grounding of phrases in images. In contrast to previous works, our model, which we call G 3 RAPHGROUND, uses graphs to formulate more complex, non-sequential dependencies among proposal image regions and phrases. We capture intra-modal dependencies using a separate graph neural network for each modality (visual and lingual), and then use conditional message-passing in another graph neural network to fuse their outputs and capture crossmodal relationships. This final representation results in grounding decisions. The framework supports many-tomany matching and is able to ground single phrase to multiple image regions and vice versa. We validate our design choices through a series of ablation studies and illustrate state-of-the-art performance on Flickr30k and ReferIt Game benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- TransVG: End-to-End Visual Grounding with TransformersJiajun Deng, Zhengyuan Yang, Tianlang Chen, Wengang Zhou 等ICCV 2021 · 被引用 468 次
- Referring Transformer: A One-step Approach to Multi-task Visual GroundingMuchen Li, Leonid SigalNeurIPS 2021 · 被引用 270 次
- Shifting More Attention to Visual Backbone: Query-modulated Refinement Networks for End-to-End Visual GroundingJiabo Ye, Junfeng Tian, Ming Yan, Xiaoshan Yang 等CVPR 2022 · 被引用 89 次
- Auto-Parsing Network for Image Captioning and Visual Question AnsweringXu Yang, Chongyang Gao, Hanwang Zhang, Jianfei CaiICCV 2021 · 被引用 45 次
- DQ-DETR: Dual Query Detection Transformer for Phrase Extraction and GroundingShilong Liu, Shijia Huang, Feng Li, Hao Zhang 等AAAI 2023 · 被引用 44 次
相关 Paper
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 被引用 100 次
- Disentangled Motif-aware Graph Learning for Phrase GroundingZongshen Mu, Siliang Tang, Jie Tan, Qiang Yu 等AAAI 2021 · 被引用 38 次
- Cross-Modal Omni Interaction Modeling for Phrase GroundingTianyu Yu, Tianrui Hui, Zhihao Yu, Yue Liao 等ACM MM 2020 · 被引用 14 次
- Visual-Semantic Graph Matching for Visual GroundingChenchen Jing, Yuwei Wu, Mingtao Pei, Yao Hu 等ACM MM 2020 · 被引用 35 次
- Improving Zero-Shot Phrase Grounding via Reasoning on External Knowledge and Spatial RelationsZhan Shi, Yilin Shen, Hongxia Jin, Xiaodan ZhuAAAI 2022 · 被引用 7 次
