MuSe: Multi-Stage Graph Reasoning via Vision-Language Models
Guanyu Wang, Xu Chu, Zhijie Tan, Xinrong Chen, Tong Mo, Weiping Li
Abstract
Graph-related tasks are traditionally addressed with Graph Neural Networks (GNNs) or graph transformers, but their task-specific training limits generalization. Large Language Models (LLMs) offer stronger generalization, yet encoding graphs as one-dimensional text struggles to capture multi-hop dependencies and two-dimensional topology. Vision-Language Models (VLMs) provide an alternative by visualizing graphs, but rendering large graphs in a single image causes clutter, occlusion, and distraction, hindering reasoning. We propose MuSe, a novel multi-stage graph reasoning framework based on VLMs. Instead of processing entire graphs at once, MuSe incrementally samples and visualizes task-relevant subgraphs, enabling progressive reasoning. The framework employs a two-stage training paradigm: supervised fine-tuning to acquire local sampling and reasoning skills, followed by reinforcement learning with GRPO to refine the sampling strategy and control dialog length. To support evaluation, we introduce LGVLQA, a new multimodal dataset with larger and more complex graph structures, addressing the scalability limitations of existing benchmarks. Experiments show that MuSe consistently out-performs leading LLM and VLM baselines, demonstrating improved structural understanding and reasoning ability. Our code and data are available at this url.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0209edcd-342a-47a7-8dab-9f74698934e3Builds on13
- LightGCN: Simplifying and Powering Graph Convolution Network for RecommendationXiangnan He, Kuan Deng, Xiang Wang, Yan Li et al.SIGIR 2020 · 4,448 citations
- MetaMath: Bootstrap Your Own Mathematical Questions for Large Language ModelsLonghui Yu, Weisen Jiang, Han Shi, Jincheng Yu et al.ICLR 2024 · 637 citations
- Learning Intents behind Interactions with Knowledge Graph for RecommendationXiang Wang, Tinglin Huang, Dingxian Wang, Yancheng Yuan et al.WWW 2021 · 584 citations
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan et al.NeurIPS 2023 · 420 citations
- Structure-Aware Transformer for Graph Representation LearningDexiong Chen, Leslie O'Bray, Karsten M. BorgwardtICML 2022 · 349 citations
Related papers
- GITA: Graph to Visual and Textual Integration for Vision-Language Graph ReasoningYanbin Wei, Shuai Fu, Weisen Jiang, Zejian Zhang et al.NeurIPS 2024 · 56 citations
- Advancement in Graph Understanding: A Multimodal Benchmark and Fine-Tuning of Vision-Language ModelsQihang Ai, Jiafan Li, Jincheng Dai, Jianwu Zhou et al.ACL 2024 · 1 citation
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal ModelsYuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue et al.CVPR 2026
- MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment GroundingFuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang et al.ACL 2026 · 10 citations
- Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningYingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang et al.ACL 2025 · 15 citations
