VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
Yunxin Li, Baotian Hu, Haoyuan Shi, Wei Wang, Longyue Wang, Min Zhang
Abstract
Large Multimodal Models (LMMs) have achieved impressive success in visual understanding and reasoning, remarkably improving the performance of mathematical reasoning in a visual context. Yet, a challenging type of visual math lies in the multimodal graph theory problem, which demands that LMMs understand the graphical structures accurately and perform multi-step reasoning on the visual graph. Additionally, exploring multimodal graph theory problems will lead to more effective strategies in fields like biology, transportation, and robotics planning. To step forward in this direction, we are the first to design a benchmark named VisionGraph, used to explore the capabilities of advanced LMMs in solving multimodal graph theory problems. It encompasses eight complex graph problem tasks, from connectivity to shortest path problems. Subsequently, we present a Description-Program-Reasoning (DPR) chain to enhance the logical accuracy of reasoning processes through graphical structure description generation and algorithm-aware multi-step reasoning. Our extensive study shows that 1) GPT-4V outperforms Gemini Pro in multi-step graph reasoning; 2) All LMMs exhibit inferior perception accuracy for graphical structures, whether in zero/few-shot settings or with supervised fine-tuning (SFT), which further affects problem-solving performance; 3) DPR significantly improves the multi-step graph reasoning capabilities of LMMs and the GPT-4V (DPR) agent achieves SOTA performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92e2e8ad-d668-4cce-8f66-2809ec92f480Cited by top-tier papers11
- GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware GenerationZhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu et al.ACM MM 2024 · 22 citations
- Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph FilteringHaoran Zheng, Renchi Yang, Hongtao Wang, Jianliang XuKDD 2026 · 7 citations
- The Underappreciated Power of Vision Models for Graph Structural UnderstandingXinjian Zhao, Wei Pang, Zhongkai Xue, Xiangru Jian et al.NeurIPS 2025 · 7 citations
- DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAsYanbin Wei, Jiangyue Yan, Chun Kang, Yang Chen et al.CVPR 2026 · 6 citations
- Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision ModelsThanh-Dat Truong, Huu-Thien Tran, Tran Thai Son, Bhiksha Raj et al.NeurIPS 2025 · 6 citations
Builds on6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li et al.ICLR 2024 · 3,079 citations
- WebVLN: Vision-and-Language Navigation on WebsitesQi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou et al.AAAI 2024 · 22 citations
- Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future DirectionsJing Gu, Eliana Stefani, Qi Wu, Jesse Thomason et al.ACL 2022
Related papers
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language ModelsChengke Zou, Xingang Guo, Rui Yang, Junyu Zhang et al.ICLR 2025
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu et al.ICLR 2026 · 20 citations
- Caption This, Reason That: VLMs Caught in the MiddleZihan Weng, Lucas Gomez, Taylor W. Webb, Pouya BashivanNeurIPS 2025 · 3 citations
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo et al.ICCV 2025 · 22 citations
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu et al.ICLR 2024 · 1,472 citations
