MLaGA: Multimodal Large Language and Graph Assistant
Dongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah, Qiaoyu Tan
摘要
Large language models (LLMs) have shown strong potential in graph learning by enabling powerful reasoning and broad generalization. However, existing Graph LLMs remain confined to textual graphs, where node features are represented solely by textual descriptions. This narrow focus overlooks a growing class of multimodal graphs (MMGs), where nodes are associated with multimodal attributes, such as text and images, commonly seen in real-world applications like e-commerce, social networks, and digital artworks. Effectively modeling MMGs presents two key challenges: (i) capturing fine-grained cross-modal interactions—e.g., between visual patches and word tokens—while preserving structural dependencies, and (ii) achieving unified generalization across diverse tasks and domains within a single model. To address these challenges, we propose MLaGA, a novel Multimodal Large Language and Graph Assistant that serves as the LLM-based foundation model for multimodal graphs. MLaGA comprises two key innovations: (1) the Structure-Aware Multimodal Aligner (SMA), which performs token-level fusion of visual and textual representations via query-driven cross-attention while explicitly maintaining graph topology, delivering high-quality multimodal node representations; and (2) the Multi-Task Multimodal Graph Instruction Tuning (MMGIT) framework, which combines structured prompts with a novel cross-task attention mechanism over task-specific projectors to enable scalable instruction tuning across multiple graph tasks, significantly improving cross-task generalization. Together, these components endow MLaGA with fine-grained multimodal reasoning and strong generalization across domains and objectives. Extensive experiments on 12 diverse MMG benchmarks demonstrate that MLaGA consistently surpasses state-of-the-art baselines on standard node classification and link prediction tasks and multimodal generative tasks, establishing a new foundation for multimodal graph learning while remaining easily extensible to new tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu 等CVPR 2026 · 被引用 2 次
- GraphVLM: Benchmarking Vision Language Models for Multimodal Graph LearningJiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha 等CVPR 2026 · 被引用 1 次
它引用的顶会 Paper15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi 等SIGIR 2024 · 被引用 182 次
相关 Paper
- LLaGA: Large Language and Graph AssistantRunjin Chen, Tong Zhao, Ajay Kumar Jaiswal, Neil Shah 等ICML 2024 · 被引用 180 次
- GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on GraphsYi Fang, Bowen Jin, Jiacheng Shen, Sirui Ding 等CVPR 2025
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 被引用 29 次
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 被引用 33 次
- GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingZiyin Zhang, Hang Yu, Sage Lee, Peng Di 等ACL 2025 · 被引用 11 次
