GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on Graphs
Yi Fang, Bowen Jin, Jiacheng Shen, Sirui Ding, Qiaoyu Tan, Jiawei Han
Abstract
The rapid development of Multimodal Large Language Models (MLLMs) has enabled the integration of multiple modalities, including texts and images, within the large language model (LLM) framework. However, texts and images are usually interconnected, forming a multimodal attributed graph (MMAG). It is underexplored how MLLMs can incorporate the relational information (i.e., graph structure) and semantic information (i.e., texts and images) on such graphs for multimodal comprehension and generation. In this paper, we propose GRAPHGPT-O, which supports omni-multimodal understanding and creation on MMAGs. We first comprehensively study linearization variants to transform semantic and structural information as input for MLLMs. Then, we propose a hierarchical aligner that enables deep graph encoding, bridging the gap between MMAGs and MLLMs. Finally, we explore the inference choices, adapting MLLM to interleaved text and image generation in graph scenarios. Extensive experiments on three datasets from different domains demonstrate the effectiveness of our proposed method.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6cb0589-c4ba-4a4f-9184-75d4d5dbc7b3Cited by top-tier papers4
- OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed GraphChenxi Wan, Xunkai Li, Yilong Zuo, Haokun Deng et al.ICML 2026 · 9 citations
- PerfGuard: A Performance-Aware Agent for Visual Content GenerationZhipeng Chen, Zhongrui Zhang, Chao Zhang, Yifan Xu et al.ICLR 2026 · 1 citation
- Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph LearningJing Zhu, Yuhang Zhou, Shengyi Qian, Zhongmou He et al.CVPR 2025
- Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based ApproachSicheng Liu, Xunkai Li, Daohan Su, Ru Zhang et al.ICML 2026
Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
Related papers
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah et al.KDD 2026 · 13 citations
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 29 citations
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi et al.ICLR 2024 · 315 citations
- Simple Yet Effective: Structure Guided Pre-trained Transformer for Multi-modal Knowledge Graph ReasoningKe Liang, Lingyuan Meng, Yue Liu, Meng Liu et al.ACM MM 2024 · 29 citations
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingHao Li, Changyao Tian, Jie Shao, Xizhou Zhu et al.CVPR 2025
