Modality-Aware Integration with Large Language Models for Knowledge-Based Visual Question Answering
Junnan Dong, Qinggang Zhang, Huachi Zhou, Daochen Zha, Pai Zheng, Xiao Huang
Abstract
Knowledge-based visual question answering (KVQA) has been extensively studied to answer visual questions with external knowledge, e.g., knowledge graphs (KGs). While several attempts have been proposed to leverage large language models (LLMs) as an implicit knowledge source, it remains challenging since LLMs may generate hallucinations. Moreover, multiple knowledge sources, e.g., images, KGs and LLMs, cannot be readily aligned for complex scenarios. To tackle these, we present a novel modality-aware integration with LLMs for KVQA (MAIL). It carefully leverages multimodal knowledge for both image understanding and knowledge reasoning. Specifically, (i) we propose a two-stage prompting strategy with LLMs to densely embody the image into a scene graph with detailed visual features; (ii) We construct a coupled concept graph by linking the mentioned entities with external facts. (iii) A tailored pseudo-siamese graph medium fusion is designed for sufficient multimodal fusion. We utilize the shared mentioned entities in two graphs as mediums to bridge a tight intermodal exchange, while maximally preserving insightful intra-modal learning by constraining the fusion within mediums. Extensive experiments show the superiority of MAIL. * Corresponding Author Recently, several studies have explored using large language models (LLMs) as supplementary knowledge bases and reasoning tools for KVQA (Yang et al., 2022; Gui et al., 2022; Lin et al., 2022); according to how they fuse the knowledge, they can be broadly categorized into direct prompting and modality-agnostic approaches, shown in Figure 1 (a) and (b), respectively. The former directly prompts the question and the corresponding image caption to LLMs for answers (Yang et al., 2022) . The latter leverages LLMs to generate candidate answers with supporting evidence and simply combines both question and the external knowledge embedding, e.g., Wikidata (Shengyuan et al., 2024), for reasoning at the final stage (Gui et al., 2022; Lin et al., 2022). While the above methods have employed LLMs in various ways for KVQA, we argue that they have not fully leveraged the knowledge from LLMs and lack the cross-modal reasoning ability, potentially resulting in sub-optimal performance for complex VQA scenarios. (i) LLMs could incorrectly answer questions or provide unreliable evidence for reasoning. On the one hand, direct prompting to LLMs may struggle to identify the right answer for many complex or domain-specific questions, due to the lack of domain knowledge (Amaro et al., 2023; Chen et al., 2024) . On the other hand, LLMs
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 482cc4f8-25b7-4752-bdc0-bc6239d6be22Cited by top-tier papers10
- Entity Alignment with Noisy Annotations from Large Language ModelsShengyuan Chen, Qinggang Zhang, Junnan Dong, Wen Hua et al.NeurIPS 2024 · 44 citations
- Youtu-GraphRAG: Vertically Unified Agents for Graph Retrieval-Augmented Complex ReasoningJunnan Dong, Siyu An, Yifei Yu, Qian-Wen Zhang et al.ICLR 2026 · 29 citations
- Tokenization, Fusion, and Augmentation: Towards Fine-grained Multi-modal Entity RepresentationYichi Zhang, Zhuo Chen, Lingbing Guo, Yajing Xu et al.AAAI 2025 · 25 citations
- Vision-and-Language Training Helps Deploy Taxonomic Knowledge but Does Not Fundamentally Alter ItYulu Qin, Dheeraj Varghese, Adam Dahlgren Lindström, Lucia Donatelli et al.NeurIPS 2025 · 11 citations
- Cost-efficient Knowledge-based Question Answering with Large Language ModelsJunnan Dong, Qinggang Zhang, Chuang Zhou, Hao Chen et al.NeurIPS 2024 · 10 citations
Builds on13
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong et al.NeurIPS 2023 · 4,013 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- An Empirical Study of GPT-3 for Few-Shot Knowledge-Based VQAZhengyuan Yang, Zhe Gan, Jianfeng Wang, Xiaowei Hu et al.AAAI 2022 · 517 citations
- Multi-Modal Answer Validation for Knowledge-Based VQAJialin Wu, Jiasen Lu, Ashish Sabharwal, Roozbeh MottaghiAAAI 2022 · 183 citations
Related papers
- KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question AnsweringZhiyang Li, Ao Ke, Yukun Cao, Xike XieACL 2026
- Notes-guided MLLM Reasoning: Enhancing MLLM with Knowledge and Visual Notes for Visual Question AnsweringWenlong Fang, Qiaofeng Wu, Jing Chen, Yun XueCVPR 2025
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- GraphVis: Boosting LLMs with Visual Knowledge Graph IntegrationYihe Deng, Chenchen Ye, Zijie Huang, Mingyu Derek Ma et al.NeurIPS 2024 · 23 citations
- Large Language Models Know What is Key Visual Entity: An LLM-assisted Multimodal Retrieval for VQAPu Jian, Donglei Yu, Jiajun ZhangEMNLP 2024 · 5 citations
