Multimodal Neural Graph Memory Networks for Visual Question Answering
Mahmoud Khademi
Abstract
We introduce a new neural network architecture, Multimodal Neural Graph Memory Networks (MN-GMN), for visual question answering. The MN-GMN uses graph structure with different region features as node attributes and applies a recently proposed powerful graph neural network model, Graph Network (GN), to reason about objects and their interactions in an image. The input module of the MN-GMN generates a set of visual features plus a set of encoded region-grounded captions (RGCs) for the image. The RGCs capture object attributes and their relationships. Two GNs are constructed from the input module using the visual features and encoded RGCs. Each node of the GNs iteratively computes a questionguided contextualized representation of the visual/textual information assigned to it. Then, to combine the information from both GNs, the nodes write the updated representations to an external spatial memory. The final states of the memory cells are fed into an answer module to predict an answer. Experiments show MN-GMN rivals the state-of-the-art models on Visual7W, VQA-v2.0, and CLEVR datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 45926c43-d3b9-424d-a6cd-71fad00ac83dCited by top-tier papers6
- Tackling Modality Heterogeneity with Multi-View Calibration Network for Multimodal Sentiment DetectionYiwei Wei, Shaozu Yuan, Ruosong Yang, Lei Shen et al.ACL 2023 · 42 citations
- Learning to Memorize Feature Hallucination for One-Shot Image GenerationYu Xie, Yanwei Fu, Ying Tai, Yun Cao et al.CVPR 2022 · 10 citations
- HybridPrompt: Bridging Language Models and Human Priors in Prompt Tuning for Visual Question AnsweringZhiyuan Ma, Zhihuan Yu, Jianjun Li, Guohui LiAAAI 2023 · 8 citations
- GHAN: Graph-Based Hierarchical Aggregation Network for Text-Video RetrievalYahan Yu, Bojie Hu, Yu LiEMNLP 2022 · 7 citations
- Modeling Temporal-Modal Entity Graph for Procedural Multimodal Machine ComprehensionHuibin Zhang, Zhengkun Zhang, Yao Zhang, Jun Wang et al.ACL 2022 · 5 citations
Builds on1
Related papers
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
- Aligned Dual Channel Graph Convolutional Network for Visual Question AnsweringQingbao Huang, Jielong Wei, Yi Cai, Changmeng Zheng et al.ACL 2020 · 79 citations
- Multiple Objects-Aware Visual Question GenerationJiayuan Xie, Yi Cai, Qingbao Huang, Tao WangACM MM 2021 · 23 citations
- Iterative Context-Aware Graph Inference for Visual DialogDan Guo, Hui Wang, Hanwang Zhang, Zheng-Jun Zha et al.CVPR 2020
- Multi-modal Attentive Graph Pooling Model for Community Question Answer MatchingJun Hu, Quan Fang, Shengsheng Qian, Changsheng XuACM MM 2020 · 10 citations
