A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal Graphs
Dongxiao He, AnKang Yang, Jitao Zhao, Di Jin
Abstract
Graph Foundation Models (GFMs) aim to learn universal patterns through large-scale pretraining on diverse graphs and generalize to open-world scenarios. While GFMs have garnered significant attention, existing works primarily focus on sigle-modal graphs. However, many real-world graphs are multimodal, consisting of structures alongside diverse features derived from modalities such as text and images. To date, exploration into Multimodal Graph Foundation Models (MGFMs) remains limited. Incorporating multimodal data provides a more comprehensive view, allowing models to learn richer semantics, thereby advancing GFMs. We are therefore motivated to explore MGFMs, where the core challenge lies in synergistically encoding structures and multimodal features to achieve effective cross-modal alignment and fusion. To this end, we propose a graph foundation model with Cross-modal Alignment and Modality-aware Expert fusion, CAME. Specifically, CAME first generates graph embeddings for each individual modality. We then introduce a multimodal multi-expert encoding mechanism, which includes a dimension-wise routing strategy to fuse multimodal information. Finally, we employ a cross-modal contrastive loss to train CAME, enabling the adaptive alignment and fusion across different modalities. Extensive experiments demonstrate the effectiveness of CAME across multiple tasks and diverse multimodal graph datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bc8d0f99-9996-4710-95d8-d5a7d3391820Builds on17
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Large-Scale Representation Learning on Graphs via BootstrappingShantanu Thakoor, Corentin Tallec, Mohammad Gheshlaghi Azar, Mehdi Azabou et al.ICLR 2022 · 311 citations
- One For All: Towards Training One Graph Model For All Classification TasksHao Liu, Jiarui Feng, Lecheng Kong, Ningyue Liang et al.ICLR 2024 · 253 citations
Related papers
- RAG-GFM: Overcoming In-Memory Bottlenecks in Graph Foundation Models via Retrieval-Augmented GenerationHaonan Yuan, Qingyun Sun, Jiacheng Tao, Xingcheng Fu et al.WWW 2026
- Toward Effective Multimodal Graph Foundation Model: A Divide-and-Conquer Based ApproachSicheng Liu, Xunkai Li, Daohan Su, Ru Zhang et al.ICML 2026
- Enhanced Expert Merging for Mixture-of-Experts in Graph Foundation ModelsLei Liu, Xingyu Xia, Qianqian Xie, Ben Liu et al.NeurIPS 2025 · 4 citations
- Graph4MM: Weaving Multimodal Learning with Structural InformationXuying Ning, Dongqi Fu, Tianxin Wei, Wujiang Xu et al.ICML 2025
- UniGraph2: Learning a Unified Embedding Space to Bind Multimodal GraphsYufei He, Yuan Sui, Xiaoxin He, Yue Liu et al.WWW 2025 · 37 citations
