Better with Less: A Data-Active Perspective on Pre-Training Graph Neural Networks
Jiarong Xu, Renhong Huang, Xin Jiang, Yuxuan Cao, Carl Yang, Chunping Wang, Yang Yang
摘要
Pre-training on graph neural networks (GNNs) aims to learn transferable knowledge for downstream tasks with unlabeled data, and it has recently become an active research area. The success of graph pre-training models is often attributed to the massive amount of input data. In this paper, however, we identify the curse of big data phenomenon in graph pre-training: more training data do not necessarily lead to better downstream performance. Motivated by this observation, we propose a better-with-less framework for graph pre-training: fewer, but carefully chosen data are fed into a GNN model to enhance pre-training. The proposed pre-training pipeline is called the data-active graph pre-training (APT) framework, and is composed of a graph selector and a pre-training model. The graph selector chooses the most representative and instructive data points based on the inherent properties of graphs as well as predictive uncertainty. The proposed predictive uncertainty, as feedback from the pre-training model, measures the confidence level of the model in the data. When fed with the chosen data, on the other hand, the pre-training model grasps an initial understanding of the new, unseen data, and at the same time attempts to remember the knowledge learned from previous data. Therefore, the integration and interaction between these two components form a unified framework (APT), in which graph pre-training is performed in a progressive and iterative way. Experiment results show that the proposed APT is able to obtain an efficient pre-training model with fewer training data and better downstream performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- GRAVER: Generative Graph Vocabularies for Robust Graph Foundation Models Fine-tuningHaonan Yuan, Qingyun Sun, Junhua Shi, Xingcheng Fu 等NeurIPS 2025 · 被引用 17 次
- Beyond Efficiency: Molecular Data Pruning for Enhanced GeneralizationDingshuo Chen, Zhixun Li, Yuyan Ni, Guibin Zhang 等NeurIPS 2024 · 被引用 14 次
- Cross-Domain Graph Data Scaling: A Showcase with Diffusion ModelsWenzhuo Tang, Haitao Mao, Danial Dervovic, Ivan Brugere 等NeurIPS 2025 · 被引用 8 次
- Tree of Preferences for Diversified RecommendationHanyang Yuan, Ning Tang, Tongya Zheng, Jiarong Xu 等NeurIPS 2025 · 被引用 3 次
- Extracting Training Data from Molecular Pre-trained ModelsRenhong Huang, Jiarong Xu, Zhiming Yang, Xiang Si 等NeurIPS 2024 · 被引用 3 次
它引用的顶会 Paper29
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong 等NeurIPS 2020 · 被引用 3,935 次
- Graph Contrastive Learning with AugmentationsYuning You, Tianlong Chen, Yongduo Sui, Ting Chen 等NeurIPS 2020 · 被引用 3,042 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- Contrastive Multi-View Representation Learning on GraphsKaveh Hassani, Amir Hosein Khas AhmadiICML 2020 · 被引用 1,663 次
- Geom-GCN: Geometric Graph Convolutional NetworksHongbin Pei, Bingzhe Wei, Kevin Chen-Chuan Chang, Yu Lei 等ICLR 2020 · 被引用 1,445 次
相关 Paper
- GPT-GNN: Generative Pre-Training of Graph Neural NetworksZiniu Hu, Yuxiao Dong, Kuansan Wang, Kai-Wei Chang 等KDD 2020 · 被引用 438 次
- Learning to Pre-train Graph Neural NetworksYuanfu Lu, Xunqiang Jiang, Yuan Fang, Chuan ShiAAAI 2021 · 被引用 158 次
- Edge Prompt Tuning for Graph Neural NetworksXingbo Fu, Yinhan He, Jundong LiICLR 2025 · 被引用 140 次
- When to Pre-Train Graph Neural Networks? From Data Generation Perspective!Yuxuan Cao, Jiarong Xu, Carl Yang, Jiaan Wang 等KDD 2023 · 被引用 18 次
- GPPT: Graph Pre-training and Prompt Tuning to Generalize Graph Neural NetworksMingchen Sun, Kaixiong Zhou, Xin He, Ying Wang 等KDD 2022 · 被引用 141 次
