Relation-Aware Graph Attention Network for Visual Question Answering
Linjie Li, Zhe Gan, Yu Cheng, Jingjing Liu
Abstract
In order to answer semantically-complicated questions about an image, a Visual Question Answering (VQA) model needs to fully understand the visual scene in the image, especially the interactive dynamics between different objects. We propose a Relation-aware Graph Attention Network (ReGAT), which encodes each image into a graph and models multi-type inter-object relations via a graph attention mechanism, to learn question-adaptive relation representations. Two types of visual object relations are explored: (i) Explicit Relations that represent geometric positions and semantic interactions between objects; and (ii) Implicit Relations that capture the hidden dynamics between image regions. Experiments demonstrate that ReGAT outperforms prior state-of-the-art approaches on both VQA 2.0 and VQA-CP v2 datasets. We further show that Re-GAT is compatible to existing VQA architectures, and can be used as a generic relation encoder to boost the model performance for VQA. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers65
- Large-Scale Adversarial Training for Vision-and-Language Representation LearningZhe Gan, Yen-Chun Chen, Linjie Li, Chen Zhu et al.NeurIPS 2020 · 561 citations
- Multi-modal Graph Fusion for Named Entity Recognition with Targeted Visual GuidanceDong Zhang, Suzhong Wei, Shoushan Li, Hanqian Wu et al.AAAI 2021 · 240 citations
- Graph Optimal Transport for Cross-Domain AlignmentLiqun Chen, Zhe Gan, Yu Cheng, Linjie Li et al.ICML 2020 · 193 citations
- Modulated Graph Convolutional Network for 3D Human Pose EstimationZhiming Zou, Wei TangICCV 2021 · 166 citations
- A Novel Graph-based Multi-modal Fusion Encoder for Neural Machine TranslationYongjing Yin, Fandong Meng, Jinsong Su, Chulun Zhou et al.ACL 2020 · 145 citations
Related papers
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
- Progressive Graph Attention Network for Video Question AnsweringLiang Peng, Shuangji Yang, Yi Bin, Guoqing WangACM MM 2021 · 47 citations
- Multimodal Neural Graph Memory Networks for Visual Question AnsweringMahmoud KhademiACL 2020 · 35 citations
- Hierarchical Graph Attention Network for Visual Relationship DetectionLi Mi, Zhenzhong ChenCVPR 2020
- Iterative Context-Aware Graph Inference for Visual DialogDan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha et al.CVPR 2020
