A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal Graphs
Dongxiao He, AnKang Yang, Jitao Zhao, Di Jin
摘要
Graph Foundation Models (GFMs) aim to learn universal patterns through large-scale pretraining on diverse graphs and generalize to open-world scenarios. While GFMs have garnered significant attention, existing works primarily focus on sigle-modal graphs. However, many real-world graphs are multimodal, consisting of structures alongside diverse features derived from modalities such as text and images. To date, exploration into Multimodal Graph Foundation Models (MGFMs) remains limited. Incorporating multimodal data provides a more comprehensive view, allowing models to learn richer semantics, thereby advancing GFMs. We are therefore motivated to explore MGFMs, where the core challenge lies in synergistically encoding structures and multimodal features to achieve effective cross-modal alignment and fusion. To this end, we propose a graph foundation model with Cross-modal Alignment and Modality-aware Expert fusion, CAME. Specifically, CAME first generates graph embeddings for each individual modality. We then introduce a multimodal multi-expert encoding mechanism, which includes a dimension-wise routing strategy to fuse multimodal information. Finally, we employ a cross-modal contrastive loss to train CAME, enabling the adaptive alignment and fusion across different modalities. Extensive experiments demonstrate the effectiveness of CAME across multiple tasks and diverse multimodal graph datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- Large-Scale Representation Learning on Graphs via BootstrappingShantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou 等ICLR 2022 · 被引用 311 次
- One For All: Towards Training One Graph Model For All Classification TasksHao Liu, Jiarui Feng, Lecheng Kong, Ningyue Liang 等ICLR 2024 · 被引用 253 次
相关 Paper
- RAG-GFM: Overcoming In-Memory Bottlenecks in Graph Foundation Models via Retrieval-Augmented GenerationHaonan Yuan, Qingyun Sun, Jiacheng Tao, Xingcheng Fu 等WWW 2026
- Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based ApproachSicheng Liu, Xunkai Li, Daohan Su, Ru Zhang 等ICML 2026
- Enhanced Expert Merging for Mixture-of-Experts in Graph Foundation ModelsLei Liu, Xingyu Xia, Qianqian Xie, Ben Liu 等NeurIPS 2025 · 被引用 4 次
- Graph4MM: Weaving Multimodal Learning with Structural InformationXuying Ning, Dongqi Fu, Tianxin Wei, Wujiang Xu 等ICML 2025
- UniGraph2: Learning a Unified Embedding Space to Bind Multimodal GraphsYufei He, Yuan Sui, Xiaoxin He, Yue Liu 等WWW 2025 · 被引用 37 次
