Boosting Visual Question Answering with Context-aware Knowledge Aggregation
Guohao Li, Xin Wang, Wenwu Zhu
Abstract
Given an image and a natural language question, Visual Question Answering (VQA) aims at answering the textual question correctly. Most VQA approaches in literature targets at finding answers to the questions solely based on analyzing the given images and questions alone. Other works that try to incorporate external knowledge into VQA adopt a query-based search on knowledge graphs to obtain the answer. However, these works suffer from the following problem: the model training process heavily relies on the ground-truth knowledge facts which serve as supervised information --- missing these ground-truth knowledge facts during training will lead to failures in producing the correct answers. To solve the challenging issue, we propose a Knowledge Graph Augmented (KG-Aug) model which conducts context-aware knowledge aggregation on external knowledge graphs, requiring no ground-truth knowledge facts for extra supervision. The proposed KG-Aug model is capable of retrieving context-aware knowledge subgraphs given visual images and textual questions, and learning to aggregate the useful image- and question-dependent knowledge which is then utilized to boost the accuracy in answering visual questions. We carry out extensive experiments to validate the effectiveness of our proposed KG-Aug models against several baseline approaches on various datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4e7f2148-cdeb-483b-a103-716039fa1309Cited by top-tier papers20
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu et al.AAAI 2022 · 517 citations
- Multi-Modal Answer Validation for Knowledge-Based VQAJialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh MottaghiAAAI 2022 · 183 citations
- REVIVE: Regional Visual Representation Matters in Knowledge-Based Visual Question AnsweringYuanze Lin, Yujia Xie, Dongdong Chen, Yichong Xu et al.NeurIPS 2022 · 119 citations
- MuKEA: Multimodal Knowledge Extraction and Accumulation for Knowledge-based Visual Question AnsweringYang Ding, Jing Yu, Bang Liu, Yue Hu et al.CVPR 2022 · 115 citations
- Fine-grained Late-interaction Multi-modal Retrieval for Retrieval Augmented Visual Question AnsweringWeizhe Lin, Jinghong Chen, Jingbiao Mei, Alexandru Coca et al.NeurIPS 2023 · 108 citations
Builds on2
- KnowIT VQA: Answering Knowledge-Based Questions about VideosNoa Garcia, Mayu Otani, Chenhui Chu, Yuta NakashimaAAAI 2020 · 93 citations
- From Strings to Things: Knowledge-Enabled VQA Model That Can Read and ReasonAjeet Kumar Singh, Anand Mishra, Shashank Shekhar, Anirban ChakrabortyICCV 2019 · 54 citations
Related papers
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- Query and Attention Augmentation for Knowledge-Based Explainable ReasoningYifeng Zhang, Ming Jiang, Qi ZhaoCVPR 2022 · 14 citations
- Hypergraph Transformer: Weakly-Supervised Multi-hop Reasoning for Knowledge-based Visual Question AnsweringYu-Jung Heo, Eun-Sol Kim, Woo Suk Choi, Byoung-Tak ZhangACL 2022
- EntRAG: Entity-Centric Retrieval-Augmented Generation for Knowledge-based Visual Question AnsweringYiheng Hu, Xiaoyang Wang, Qing Liu, Sherry Xu et al.ICML 2026
- Dynamic Key-Value Memory Enhanced Multi-Step Graph Reasoning for Knowledge-Based Visual Question AnsweringMingxiao Li, Marie-Francine MoensAAAI 2022 · 20 citations
