Lune

KDD2026Top-tier venue

MLaGA: Multimodal Large Language and Graph Assistant

Dongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah, Qiaoyu Tan

2026Year
13Citations
2Top-tier citations

Abstract

Large language models (LLMs) have shown strong potential in graph learning by enabling powerful reasoning and broad generalization. However, existing Graph LLMs remain confined to textual graphs, where node features are represented solely by textual descriptions. This narrow focus overlooks a growing class of multimodal graphs (MMGs), where nodes are associated with multimodal attributes, such as text and images, commonly seen in real-world applications like e-commerce, social networks, and digital artworks. Effectively modeling MMGs presents two key challenges: (i) capturing fine-grained cross-modal interactions—e.g., between visual patches and word tokens—while preserving structural dependencies, and (ii) achieving unified generalization across diverse tasks and domains within a single model. To address these challenges, we propose MLaGA, a novel Multimodal Large Language and Graph Assistant that serves as the LLM-based foundation model for multimodal graphs. MLaGA comprises two key innovations: (1) the Structure-Aware Multimodal Aligner (SMA), which performs token-level fusion of visual and textual representations via query-driven cross-attention while explicitly maintaining graph topology, delivering high-quality multimodal node representations; and (2) the Multi-Task Multimodal Graph Instruction Tuning (MMGIT) framework, which combines structured prompts with a novel cross-task attention mechanism over task-specific projectors to enable scalable instruction tuning across multiple graph tasks, significantly improving cross-task generalization. Together, these components endow MLaGA with fine-grained multimodal reasoning and strong generalization across domains and objectives. Extensive experiments on 12 diverse MMG benchmarks demonstrate that MLaGA consistently surpasses state-of-the-art baselines on standard node classification and link prediction tasks and multimodal generative tasks, establishing a new foundation for multimodal graph learning while remaining easily extensible to new tasks.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 42b59491-5a37-4412-9353-b2b8dbdeb895

Cited by top-tier papers2

Ask how each one uses it

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines