Lune

KDD2026顶会

MLaGA: Multimodal Large Language and Graph Assistant

Dongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah, Qiaoyu Tan

2026年份
13被引次数
2顶会引用

摘要

Large language models (LLMs) have shown strong potential in graph learning by enabling powerful reasoning and broad generalization. However, existing Graph LLMs remain confined to textual graphs, where node features are represented solely by textual descriptions. This narrow focus overlooks a growing class of multimodal graphs (MMGs), where nodes are associated with multimodal attributes, such as text and images, commonly seen in real-world applications like e-commerce, social networks, and digital artworks. Effectively modeling MMGs presents two key challenges: (i) capturing fine-grained cross-modal interactions—e.g., between visual patches and word tokens—while preserving structural dependencies, and (ii) achieving unified generalization across diverse tasks and domains within a single model. To address these challenges, we propose MLaGA, a novel Multimodal Large Language and Graph Assistant that serves as the LLM-based foundation model for multimodal graphs. MLaGA comprises two key innovations: (1) the Structure-Aware Multimodal Aligner (SMA), which performs token-level fusion of visual and textual representations via query-driven cross-attention while explicitly maintaining graph topology, delivering high-quality multimodal node representations; and (2) the Multi-Task Multimodal Graph Instruction Tuning (MMGIT) framework, which combines structured prompts with a novel cross-task attention mechanism over task-specific projectors to enable scalable instruction tuning across multiple graph tasks, significantly improving cross-task generalization. Together, these components endow MLaGA with fine-grained multimodal reasoning and strong generalization across domains and objectives. Extensive experiments on 12 diverse MMG benchmarks demonstrate that MLaGA consistently surpasses state-of-the-art baselines on standard node classification and link prediction tasks and multimodal generative tasks, establishing a new foundation for multimodal graph learning while remaining easily extensible to new tasks.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper15

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖