Lune

CVPR2026顶会

VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal Models

Yuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue, Yijing Chen, Yue Fan, Bo Zhang, Qian Li, Lizhen Cui

出版方
2026年份

摘要

Understanding and reasoning over structured knowledge is a fundamental capability for intelligent systems. While Large Language Models (LLMs) have leveraged textual Knowledge Graphs (KGs) for relational reasoning, linearizing graph structures into text leads to loss of higher-order relational cues. Inspired by the advances of Large Multimodal Models (LMMs) to capture higher-order relational structures, we propose a novel paradigm of visualized knowledge representation, where KGs are transformed into graphical visualizations that LMMs can directly perceive and reason over. To systematically evaluate this capability, we introduce VKG-QA, a benchmark for Visual KG-based Question Answering covering three major categories and fourteen subtasks. VKG-QA is constructed via a semi-automatic pipeline ensuring high-quality, semantically aligned, and visually clear data. We evaluate 19 representative LMMs on VKG-QA and perform extensive quantitative and qualitative analyses. Results reveal that current models struggle with visualized relational understanding, graph-specific comprehension remains challenging, and closed-source models significantly outperform open-source counterparts. This thus highlights critical limitations in current LMMs and provides a scalable platform for advancing graph-aware visual reasoning.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext f2d6aff9-a8fb-4ddc-a16f-d4c3fcee8bee

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖