Better with Less: A Data-Active Perspective on Pre-Training Graph Neural Networks
Jiarong Xu, Renhong Huang, Xin Jiang, Yuxuan Cao, Carl Yang, Chunping Wang, Yang Yang
Abstract
Pre-training on graph neural networks (GNNs) aims to learn transferable knowledge for downstream tasks with unlabeled data, and it has recently become an active research area. The success of graph pre-training models is often attributed to the massive amount of input data. In this paper, however, we identify the curse of big data phenomenon in graph pre-training: more training data do not necessarily lead to better downstream performance. Motivated by this observation, we propose a better-with-less framework for graph pre-training: fewer, but carefully chosen data are fed into a GNN model to enhance pre-training. The proposed pre-training pipeline is called the data-active graph pre-training (APT) framework, and is composed of a graph selector and a pre-training model. The graph selector chooses the most representative and instructive data points based on the inherent properties of graphs as well as predictive uncertainty. The proposed predictive uncertainty, as feedback from the pre-training model, measures the confidence level of the model in the data. When fed with the chosen data, on the other hand, the pre-training model grasps an initial understanding of the new, unseen data, and at the same time attempts to remember the knowledge learned from previous data. Therefore, the integration and interaction between these two components form a unified framework (APT), in which graph pre-training is performed in a progressive and iterative way. Experiment results show that the proposed APT is able to obtain an efficient pre-training model with fewer training data and better downstream performance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 871c0fd9-7ac8-48e7-9334-88e23c0ef0b3Cited by top-tier papers11
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuningHaonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu et al.NeurIPS 2025 · 17 citations
- Beyond Efficiency: Molecular Data Pruning for Enhanced GeneralizationDingshuo Chen, Zhixun Li, Yuyan Ni, Guibin Zhang et al.NeurIPS 2024 · 14 citations
- Cross-Domain Graph Data Scaling: A Showcase with Diffusion ModelsWenzhuo Tang, Haitao Mao, Danial Dervovic, Ivan Brugere et al.NeurIPS 2025 · 8 citations
- Tree of Preferences for Diversified RecommendationHanyang Yuan, Ning Tang, Tongya Zheng, Jiarong Xu et al.NeurIPS 2025 · 3 citations
- Extracting Training Data from Molecular Pre-trained ModelsRenhong Huang, Jiarong Xu, Zhiming Yang, Xiang Si et al.NeurIPS 2024 · 3 citations
Builds on29
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen et al.NeurIPS 2020 · 3,042 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Contrastive Multi-View Representation Learning on GraphsKaveh Hassani, Amir Hosein Khas AhmadiICML 2020 · 1,663 citations
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei et al.ICLR 2020 · 1,445 citations
Related papers
- GPT-GNN: Generative Pre-Training of Graph Neural NetworksZiniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang et al.KDD 2020 · 438 citations
- Learning to Pre-train Graph Neural NetworksYuanfu Lu, Xunqiang Jiang, Yuan Fang, Chuan ShiAAAI 2021 · 158 citations
- Edge Prompt Tuning for Graph Neural NetworksXingbo Fu, Yinhan He, Jundong LiICLR 2025 · 140 citations
- When to Pre-Train Graph Neural Networks? From Data Generation Perspective!Yuxuan Cao, Jiarong Xu, Carl Yang, Jiaan Wang et al.KDD 2023 · 18 citations
- GPPT: Graph Pre-training and Prompt Tuning to Generalize Graph Neural NetworksMingchen Sun, Kaixiong Zhou, Xin He, Ying Wang et al.KDD 2022 · 141 citations
