GITA: Graph to Visual and Textual Integration for Vision-Language Graph Reasoning
Yanbin Wei, Shuai Fu, Weisen Jiang, Zejian Zhang, Zhixiong Zeng, Qi Wu, James T. Kwok, Yu Zhang
Abstract
Large Language Models (LLMs) are increasingly used for various tasks with graph structures. Though LLMs can process graph information in a textual format, they overlook the rich vision modality, which is an intuitive way for humans to comprehend structural information and conduct general graph reasoning. The potential benefits and capabilities of representing graph structures as visual images (i.e., ) are still unexplored. To fill the gap, we innovatively propose an end-to-end framework, called raph to vsual and extual Integrtion (GITA), which firstly incorporates visual graphs into general graph reasoning. Besides, we establish raph-based ision-anguage uestion nswering (GVLQA) dataset from existing graph data, which is the first vision-language dataset for general graph reasoning purposes. Extensive experiments on the GVLQA dataset and five real-world datasets show that GITA outperforms mainstream LLMs in terms of general graph reasoning capabilities. Moreover, We highlight the effectiveness of the layout augmentation on visual graphs and pretraining on the GVLQA dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 54c5a560-18a0-4089-9164-085b26c1cc01Cited by top-tier papers22
- RouterDC: Query-Based Router by Dual Contrastive Learning for Assembling Large Language ModelsShuhao Chen, Weisen Jiang, Baijiong Lin, James T. Kwok et al.NeurIPS 2024 · 113 citations
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsHuajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen et al.NeurIPS 2025 · 45 citations
- NAUTILUS: A Large Multimodal Model for Underwater Scene UnderstandingWei Xu, Cheng Wang, Dingkang Liang, Zongchuang Zhao et al.NeurIPS 2025 · 16 citations
- Graph is a Substrate Across Data ModalitiesZiming Li, Xiao-Ming Wu, Zehong Wang, Jiazheng Li et al.ICML 2026 · 16 citations
- Benchmarking and Improving Large Vision-Language Models for Fundamental Visual Graph Understanding and ReasoningYingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang et al.ACL 2025 · 15 citations
Builds on23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- GraphVis: Boosting LLMs with Visual Knowledge Graph IntegrationYihe Deng, Chenchen Ye, Zijie Huang, Mingyu Derek Ma et al.NeurIPS 2024 · 23 citations
- MuSe: Multi-Stage Graph Reasoning via Vision-Language ModelsGuanyu Wang, Xu Chu, Zhijie Tan, Xinrong Chen et al.ACL 2026
- Advancement in Graph Understanding: A Multimodal Benchmark and Fine-Tuning of Vision-Language ModelsQihang Ai, Jiafan Li, Jincheng Dai, Jianwu Zhou et al.ACL 2024 · 1 citation
- VKG-QA: Visual Knowledge Graph-based Question Answer for Large Multimodal ModelsYuntao Du, Yiming Wang, Renshuo Yuan, Jincheng Yue et al.CVPR 2026
- Visual Graph Arena: Evaluating Visual Conceptualization of Vision and Multimodal Large Language ModelsZahra Babaiee, Peyman M. Kiasari, Daniela Rus, Radu GrosuICML 2025
