Advancing Graph Foundation Models: A Data-Centric Perspective
Yuhan Li, Yuyao Wang, Jianheng Tang, Heng Chang, Yuxiang Ren, Jia Li
Abstract
Recently, Graph Foundation Models (GFMs) have emerged as a significant research topic in graph machine learning. Compared with traditional graph neural networks, GFMs demonstrate impressive zero-shot generalization across different domains and tasks through large-scale pre-training on extensive and diverse graph data. Despite the initial success of pre-training, existing GFMs face challenges such as extreme time consumption and the presence of redundancy and noise in pre-training data. To alleviate these issues, we present the first exploration of data-centric GFM, which aims to optimize pre-training data (i.e., a set of subgraphs) to establish a more efficient GFM while maintaining robust performance across various downstream tasks. We propose DCGFM, a plug-and-play approach for Data-Centric GFM that incorporates the idea of data pruning to remove redundant and less informative subgraphs from the pre-training data, thereby improving both efficiency and effectiveness. Specifically, DCGFM consists of two components: (1) a model-agnostic hard pruning module that filters out subgraphs with lower informativity scores by considering both the semantics and structures of subgraphs; and (2) a model-aware soft pruning module that dynamically prunes subgraphs with lower loss values in each pre-training epoch with a gradient rescaling strategy. Extensive experiments on representative GFM backbones demonstrate DCGFM's efficiency and effectiveness. Remarkably, DCGFM achieves even better performance using only 30% of the pre-training data. Codes and data are available at https://github.com/Yuhan1i/DCGFM.
Ask about this paper
Ask your agent about it.
Lune has read the top-tier papers around this one, so every answer names the papers it rests on.
Your agent calls
Lunesearch_papers
Free to start. No credit card required.
Terminal
Install the CLIlune papers get b54be0d3-2b1f-4f15-bcc4-f1aa0af5bbe8Cited by top-tier papers1
Ask how each one uses itRelated papers
- Towards Effective Federated Graph Foundation Model via Mitigating Knowledge EntanglementYinlin Zhu, Xunkai Li, Jishuo Jia, Miao Hu et al.NeurIPS 2025 · 17 citations
- AnomalyGFM: Graph Foundation Model for Zero/Few-shot Anomaly DetectionHezhe Qiao, Chaoxi Niu, Ling Chen, Guansong PangKDD 2025 · 8 citations
- Towards Graph Foundation Models: Learning Generalities Across Graphs via Task-TreesZehong Wang, Zheyuan Zhang, Tianyi Ma, Nitesh V. Chawla et al.ICML 2025
- GFMate: Empowering Graph Foundation Models with Test-time Prompt TuningYan Jiang, Ruihong Qiu, Zi HuangICML 2026
- GFT: Graph Foundation Model with Transferable Tree VocabularyZehong Wang, Zheyuan Zhang, Nitesh V. Chawla, Chuxu Zhang et al.NeurIPS 2024 · 108 citations
