MLaGA: Multimodal Large Language and Graph Assistant
Dongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah, Qiaoyu Tan
Abstract
Large language models (LLMs) have shown strong potential in graph learning by enabling powerful reasoning and broad generalization. However, existing Graph LLMs remain confined to textual graphs, where node features are represented solely by textual descriptions. This narrow focus overlooks a growing class of multimodal graphs (MMGs), where nodes are associated with multimodal attributes, such as text and images, commonly seen in real-world applications like e-commerce, social networks, and digital artworks. Effectively modeling MMGs presents two key challenges: (i) capturing fine-grained cross-modal interactions—e.g., between visual patches and word tokens—while preserving structural dependencies, and (ii) achieving unified generalization across diverse tasks and domains within a single model. To address these challenges, we propose MLaGA, a novel Multimodal Large Language and Graph Assistant that serves as the LLM-based foundation model for multimodal graphs. MLaGA comprises two key innovations: (1) the Structure-Aware Multimodal Aligner (SMA), which performs token-level fusion of visual and textual representations via query-driven cross-attention while explicitly maintaining graph topology, delivering high-quality multimodal node representations; and (2) the Multi-Task Multimodal Graph Instruction Tuning (MMGIT) framework, which combines structured prompts with a novel cross-task attention mechanism over task-specific projectors to enable scalable instruction tuning across multiple graph tasks, significantly improving cross-task generalization. Together, these components endow MLaGA with fine-grained multimodal reasoning and strong generalization across domains and objectives. Extensive experiments on 12 diverse MMG benchmarks demonstrate that MLaGA consistently surpasses state-of-the-art baselines on standard node classification and link prediction tasks and multimodal generative tasks, establishing a new foundation for multimodal graph learning while remaining easily extensible to new tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 42b59491-5a37-4412-9353-b2b8dbdeb895Cited by top-tier papers2
- Mario: Multimodal Graph Reasoning with Large Language ModelsYuanfu Sun, Kang Li, Pengkang Guo, Jiajin Liu et al.CVPR 2026 · 2 citations
- GraphVLM: Benchmarking Vision Language Models for Multimodal Graph LearningJiajin Liu, Dongzhe Fan, Chuanhao Ji, Daochen Zha et al.CVPR 2026 · 1 citation
Builds on15
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 7,873 citations
- GraphGPT: Graph Instruction Tuning for Large Language ModelsJiabin Tang, Yuhao Yang, Wei Wei, Lei Shi et al.SIGIR 2024 · 182 citations
Related papers
- LLaGA: Large Language and Graph AssistantRunjin Chen, Tong Zhao, Ajay Kumar Jaiswal, Neil Shah et al.ICML 2024 · 180 citations
- GRAPHGPT-O: Synergistic Multimodal Comprehension and Generation on GraphsYi Fang, Bowen Jin, Jiacheng Shen, Sirui Ding et al.CVPR 2025
- Multimodal Reasoning with Multimodal Knowledge GraphJunlin Lee, Yequan Wang, Jing Li, Min ZhangACL 2024 · 29 citations
- LLaMo: Large Language Model-based Molecular Graph AssistantJinyoung Park, Minseong Bae, Dohwan Ko, Hyunwoo J. KimNeurIPS 2024 · 33 citations
- GALLa: Graph Aligned Large Language Models for Improved Source Code UnderstandingZiyin Zhang, Hang Yu, Sage Lee, Peng Di et al.ACL 2025 · 11 citations
