GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on Graphs
Yi Fang, Bowen Jin, Jiacheng Shen, Sirui Ding, Qiaoyu Tan, Jiawei Han
摘要
The rapid development of Multimodal Large Language Models (MLLMs) has enabled the integration of multiple modalities, including texts and images, within the large language model (LLM) framework. However, texts and images are usually interconnected, forming a multimodal attributed graph (MMAG). It is underexplored how MLLMs can incorporate the relational information (i.e., graph structure) and semantic information (i.e., texts and images) on such graphs for multimodal comprehension and generation. In this paper, we propose GRAPHGPT-O, which supports omni-multimodal understanding and creation on MMAGs. We first comprehensively study linearization variants to transform semantic and structural information as input for MLLMs. Then, we propose a hierarchical aligner that enables deep graph encoding, bridging the gap between MMAGs and MLLMs. Finally, we explore the inference choices, adapting MLLM to interleaved text and image generation in graph scenarios. Extensive experiments on three datasets from different domains demonstrate the effectiveness of our proposed method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed GraphChenxi Wan, Xunkai Li, Yilong Zuo, Haokun Deng 等ICML 2026 · 被引用 9 次
- PerfGuard: A Performance-Aware Agent for Visual Content GenerationZhipeng Chen, Zhongrui Zhang, Chao Zhang, Yifan Xu 等ICLR 2026 · 被引用 1 次
- Mosaic of Modalities: A Comprehensive Benchmark for Multimodal Graph LearningJing Zhu, Yuhang Zhou, Shengyi Qian, Zhongmou He 等CVPR 2025
- Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based ApproachSicheng Liu, Xunkai Li, Daohan Su, Ru Zhang 等ICML 2026
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
相关 Paper
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah 等KDD 2026 · 被引用 13 次
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 被引用 29 次
- DreamLLM: Synergistic Multimodal Comprehension and CreationRunpei Dong, Chunrui Han, Yuang Peng, Zekun Qi 等ICLR 2024 · 被引用 315 次
- Simple Yet Effective: Structure Guided Pre-trained Transformer for Multi-modal Knowledge Graph ReasoningKe Liang, Lingyuan Meng, Yue Liu, Meng Liu 等ACM MM 2024 · 被引用 29 次
- SynerGen-VL: Towards Synergistic Image Understanding and Generation with Vision Experts and Token FoldingHao Li, Changyao Tian, Jie Shao, Xizhou Zhu 等CVPR 2025
