VisionGraph: Leveraging Large Multimodal Models for Graph Theory Problems in Visual Context
Yunxin Li, Baotian Hu, Haoyuan Shi, Wei Wang, Longyue Wang, Min Zhang
摘要
Large Multimodal Models (LMMs) have achieved impressive success in visual understanding and reasoning, remarkably improving the performance of mathematical reasoning in a visual context. Yet, a challenging type of visual math lies in the multimodal graph theory problem, which demands that LMMs understand the graphical structures accurately and perform multi-step reasoning on the visual graph. Additionally, exploring multimodal graph theory problems will lead to more effective strategies in fields like biology, transportation, and robotics planning. To step forward in this direction, we are the first to design a benchmark named VisionGraph, used to explore the capabilities of advanced LMMs in solving multimodal graph theory problems. It encompasses eight complex graph problem tasks, from connectivity to shortest path problems. Subsequently, we present a Description-Program-Reasoning (DPR) chain to enhance the logical accuracy of reasoning processes through graphical structure description generation and algorithm-aware multi-step reasoning. Our extensive study shows that 1) GPT-4V outperforms Gemini Pro in multi-step graph reasoning; 2) All LMMs exhibit inferior perception accuracy for graphical structures, whether in zero/few-shot settings or with supervised fine-tuning (SFT), which further affects problem-solving performance; 3) DPR significantly improves the multi-step graph reasoning capabilities of LMMs and the GPT-4V (DPR) agent achieves SOTA performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- GPT4Video: A Unified Multimodal Large Language Model for lnstruction-Followed Understanding and Safety-Aware GenerationZhanyu Wang, Longyue Wang, Zhen Zhao, Minghao Wu 等ACM MM 2024 · 被引用 22 次
- Cross-Contrastive Clustering for Multimodal Attributed Graphs with Dual Graph FilteringHaoran Zheng, Renchi Yang, Hongtao Wang, Jianliang XuKDD 2026 · 被引用 7 次
- The Underappreciated Power of Vision Models for Graph Structural UnderstandingXinjian Zhao, Wei Pang, Zhongkai Xue, Xiangru Jian 等NeurIPS 2025 · 被引用 7 次
- DynamicGTR: Leveraging Graph Topology Representation Preferences to Boost VLM Capabilities on Graph QAsYanbin Wei, Jiangyue Yan, Chun Kang, Yang Chen 等CVPR 2026 · 被引用 6 次
- Directed-Tokens: A Robust Multi-Modality Alignment Approach to Large Language-Vision ModelsThanh-Dat Truong, Huu-Thien Tran, Tran Thai Son, Bhiksha Raj 等NeurIPS 2025 · 被引用 6 次
它引用的顶会 Paper6
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language ModelsDeyao Zhu, Jun Chen, Xiaoqian Shen, Xiang Li 等ICLR 2024 · 被引用 3,079 次
- WebVLN: Vision-and-Language Navigation on WebsitesQi Chen, Dileepa Pitawela, Chongyang Zhao, Gengze Zhou 等AAAI 2024 · 被引用 22 次
- Vision-and-Language Navigation: A Survey of Tasks, Methods, and Future DirectionsJing Gu, Eliana Stefani, Qi Wu, Jesse Thomason 等ACL 2022
相关 Paper
- DynaMath: A Dynamic Visual Benchmark for Evaluating Mathematical Reasoning Robustness of Vision Language ModelsChengke Zou, Xingang Guo, Rui Yang, Junyu Zhang 等ICLR 2025
- VideoReasonBench: Can MLLMs Perform Vision-Centric Complex Video Reasoning?Yuanxin Liu, Kun Ouyang, Haoning Wu, Yi Liu 等ICLR 2026 · 被引用 20 次
- Caption This, Reason That: VLMs Caught in the MiddleZihan Weng, Lucas Gomez, Taylor W. Webb, Pouya BashivanNeurIPS 2025 · 被引用 3 次
- VLRMBench: A Comprehensive and Challenging Benchmark for Vision-Language Reward ModelsJiacheng Ruan, Wenzhen Yuan, Xiqi Gao, Ye Guo 等ICCV 2025 · 被引用 22 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
