Lune

CVPR2026Top-tier venue

VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models

Yuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue, Yijing Chen, Yue Fan, Bo Zhang, Qian Li, Lizhen Cui

2026Year

Abstract

Understanding and reasoning over structured knowledge is a fundamental capability for intelligent systems. While Large Language Models (LLMs) have leveraged textual Knowledge Graphs (KGs) for relational reasoning, linearizing graph structures into text leads to loss of higher-order relational cues. Inspired by the advances of Large Multimodal Models (LMMs) to capture higher-order relational structures, we propose a novel paradigm of visualized knowledge representation, where KGs are transformed into graphical visualizations that LMMs can directly perceive and reason over. To systematically evaluate this capability, we introduce VKG-QA, a benchmark for Visual KG-based Question Answering covering three major categories and fourteen subtasks. VKG-QA is constructed via a semi-automatic pipeline ensuring high-quality, semantically aligned, and visually clear data. We evaluate 19 representative LMMs on VKG-QA and perform extensive quantitative and qualitative analyses. Results reveal that current models struggle with visualized relational understanding, graph-specific comprehension remains challenging, and closed-source models significantly outperform open-source counterparts. This thus highlights critical limitations in current LMMs and provides a scalable platform for advancing graph-aware visual reasoning.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext f2d6aff9-a8fb-4ddc-a16f-d4c3fcee8bee

Builds on11

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines