VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models
Yuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue, Yijing Chen, Yue Fan, Bo Zhang, Qian Li, Lizhen Cui
Abstract
Understanding and reasoning over structured knowledge is a fundamental capability for intelligent systems. While Large Language Models (LLMs) have leveraged textual Knowledge Graphs (KGs) for relational reasoning, linearizing graph structures into text leads to loss of higher-order relational cues. Inspired by the advances of Large Multimodal Models (LMMs) to capture higher-order relational structures, we propose a novel paradigm of visualized knowledge representation, where KGs are transformed into graphical visualizations that LMMs can directly perceive and reason over. To systematically evaluate this capability, we introduce VKG-QA, a benchmark for Visual KG-based Question Answering covering three major categories and fourteen subtasks. VKG-QA is constructed via a semi-automatic pipeline ensuring high-quality, semantically aligned, and visually clear data. We evaluate 19 representative LMMs on VKG-QA and perform extensive quantitative and qualitative analyses. Results reveal that current models struggle with visualized relational understanding, graph-specific comprehension remains challenging, and closed-source models significantly outperform open-source counterparts. This thus highlights critical limitations in current LMMs and provides a scalable platform for advancing graph-aware visual reasoning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f2d6aff9-a8fb-4ddc-a16f-d4c3fcee8beeBuilds on11
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
- MM-Vet: Evaluating Large Multimodal Models for Integrated CapabilitiesWeihao Yu, Zhengyuan Yang, Linjie Li, Jianfeng Wang et al.ICML 2024 · 1,191 citations
Related papers
- M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao et al.ACL 2026
- GraphVis: Boosting LLMs with Visual Knowledge Graph IntegrationYihe Deng, Chenchen Ye, Zijie Huang, Mingyu Derek Ma et al.NeurIPS 2024 · 23 citations
- mKG-RAG: Leveraging Multimodal Knowledge Graphs in Retrieval-Augmented Generation for Knowledge-intensive VQAXu Yuan, Liangbo Ning, Qingqing Ye, Wenqi Fan et al.SIGIR 2026 · 2 citations
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 29 citations
- GraphVLM: Benchmarking Vision Language Models for Multimodal Graph LearningJiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha et al.CVPR 2026 · 1 citation
