Dynamic Key-Value Memory Enhanced Multi-Step Graph Reasoning for Knowledge-Based Visual Question Answering
Mingxiao Li, Marie-Francine Moens
Abstract
Knowledge-based visual question answering (VQA) is a vision-language task that requires an agent to correctly answer image-related questions using knowledge that is not presented in the given image. It is not only a more challenging task than regular VQA but also a vital step towards building a general VQA system. Most existing knowledge-based VQA systems process knowledge and image information similarly and ignore the fact that the knowledge base (KB) contains complete information about a triplet, while the extracted image information might be incomplete as the relations between two objects are missing or wrongly detected. In this paper, we propose a novel model named dynamic knowledge memory enhanced multi-step graph reasoning (DMMGR), which performs explicit and implicit reasoning over a key-value knowledge memory module and a spatial-aware image graph, respectively. Specifically, the memory module learns a dynamic knowledge representation and generates a knowledge-aware question representation at each reasoning step. Then, this representation is used to guide a graph attention operator over the spatial-aware image graph. Our model achieves new stateof-the-art accuracy on the KRVQR and FVQA datasets. We also conduct ablation experiments to prove the effectiveness of each component of the proposed model. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 82560fbf-bdb9-4f7c-943c-bae44c016785Cited by top-tier papers4
- VQA-GNN: Reasoning with Multimodal Knowledge via Graph Neural Networks for Visual Question AnsweringYanan Wang, Michihiro Yasunaga, Hongyu Ren, Shinya Wada et al.ICCV 2023 · 42 citations
- Improving Zero-shot Visual Question Answering via Large Language Models with Reasoning Question PromptsYunshi Lan, Xiang Li, Xin Liu, Yang Li et al.ACM MM 2023 · 29 citations
- KALM: Knowledge-Aware Integration of Local, Document, and Global Contexts for Long Document UnderstandingShangbin Feng, Zhaoxuan Tan, Wenqian Zhang, Zhenyu Lei et al.ACL 2023 · 5 citations
- Object Attribute Matters in Visual Question AnsweringPeize Li, Qingyi Si, Peng Fu, Zheng Lin et al.AAAI 2024 · 1 citation
Builds on1
Related papers
- Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question AnsweringYu-Jung Heo, Eun-Sol Kim, Woo Suk Choi, Byoung-Tak ZhangACL 2022
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringWenlong Fang, Qiaofeng Wu, Jing Chen, Yun XueCVPR 2025
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- Reveal: Retrieval-Augmented Visual-Language Pre-Training with Multi-Source Multimodal Knowledge MemoryZiniu Hu, Ahmet Iscen, Chen Sun, Zirui Wang et al.CVPR 2023
- Query and Attention Augmentation for Knowledge-Based Explainable ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2022 · 14 citations
