UniGraph2: Learning a Unified Embedding Space to Bind Multimodal Graphs
Yufei He, Yuan Sui, Xiaoxin He, Yue Liu, Yifei Sun, Bryan Hooi
Abstract
Existing foundation models, such as CLIP, aim to learn a unified embedding space for multimodal data, enabling a wide range of downstream web-based applications like search, recommendation, and content classification. However, these models often overlook the inherent graph structures in multimodal datasets, where entities and their relationships are crucial. Multimodal graphs (MMGs) represent such graphs where each node is associated with features from different modalities, while the edges capture the relationships between these entities. On the other hand, existing graph foundation models primarily focus on text-attributed graphs (TAGs) and are not designed to handle the complexities of MMGs. To address these limitations, we propose UniGraph2 1 , a novel cross-domain graph foundation model that enables general representation learning on MMGs, providing a unified embedding space. UniGraph2 employs modality-specific encoders alongside a graph neural network (GNN) to learn a unified low-dimensional embedding space that captures both the multimodal information and the underlying graph structure. We propose a new cross-domain multi-graph pre-training algorithm at scale to ensure effective transfer learning across diverse graph domains and modalities. Additionally, we adopt a Mixture of Experts (MoE) component to align features from different domains and modalities, ensuring coherent and robust embeddings that unify the information across modalities. Extensive experiments on a variety of multimodal graph tasks demonstrate that UniGraph2 significantly outperforms state-of-the-art models in tasks such as representation learning, transfer learning, and multimodal generative tasks, offering a scalable and flexible solution for learning on MMGs. CCS Concepts • Information systems → Data mining; Social networks; • Computing methodologies → Neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 304100b9-d761-47fd-9e26-a9b47001f431Cited by top-tier papers15
- EvoTest: Evolutionary Test-Time Learning for Self-Improving Agentic SystemsYufei He, Juncheng Liu, Yue Liu, Yibo Li et al.ICLR 2026 · 36 citations
- Can Knowledge Graphs Make Large Language Models More Trustworthy? An Empirical Study Over Open-ended Question AnsweringYuan Sui, Yufei He, Zifeng Ding, Bryan HooiACL 2025 · 29 citations
- Towards Effective Federated Graph Foundation Model via Mitigating Knowledge EntanglementYinlin Zhu, Xunkai Li, Jishuo Jia, Miao Hu et al.NeurIPS 2025 · 17 citations
- MLaGA: Multimodal Large Language and Graph AssistantDongzhe Fan, Jiajin Liu, Yi Fang, Djellel Difallah et al.KDD 2026 · 13 citations
- OpenMAG: A Comprehensive Benchmark for Multimodal-Attributed GraphChenxi Wan, Xunkai Li, Yilong Zuo, Haokun Deng et al.ICML 2026 · 9 citations
Builds on26
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
Related papers
- UniGraph: Learning a Unified Cross-Domain Foundation Model for Text-Attributed GraphsYufei He, Yuan Sui, Xiaoxin He, Bryan HooiKDD 2025 · 8 citations
- A Graph Foundation Model with Cross-Modal Alignment and Modality-Aware Expert Fusion for Multi-Modal GraphsDongxiao He, AnKang Yang, Jitao Zhao, Di JinICML 2026
- Can Graph Neural Networks Learn Language with Extremely Weak Text Supervision?Zihao Li, Lecheng Zheng, Bowen Jin, Dongqi Fu et al.ACL 2025
- MUG: Meta-path-aware Universal Heterogeneous Graph Pre-TrainingLianze Shan, Jitao Zhao, Dongxiao He, Yongqi Huang et al.AAAI 2026 · 1 citation
- Multi-Domain Generalized Graph Meta LearningMingkai Lin, Wenzhong Li, Ding Li, Yizhou Chen et al.AAAI 2023 · 19 citations
