Confidence May Cheat: Self-Training on Graph Neural Networks under Distribution Shift
Hongrui Liu, Binbin Hu, Xiao Wang, Chuan Shi, Zhiqiang Zhang, Jun Zhou
Abstract
Graph Convolutional Networks (GCNs) have recently attracted vast interest and achieved state-of-the-art performance on graphs, but its success could typically hinge on careful training with amounts of expensive and time-consuming labeled data. To alleviate labeled data scarcity, self-training methods have been widely adopted on graphs by labeling high-confidence unlabeled nodes and then adding them to the training step. In this line, we empirically make a thorough study for current self-training methods on graphs. Surprisingly, we find that high-confidence unlabeled nodes are not always useful, and even introduce the distribution shift issue between the original labeled dataset and the augmented dataset by self-training, severely hindering the capability of self-training on graphs. To this end, in this paper, we propose a novel Distribution Recovered Graph Self-Training framework (DR-GST), which could recover the distribution of the original labeled dataset. Specifically, we first prove the equality of loss function in self-training framework under the distribution shift case and the population distribution if each pseudo-labeled node is weighted by a proper coefficient. Considering the intractability of the coefficient, we then propose to replace the coefficient with the information gain after observing the same changing trend between them, where information gain is respectively estimated via both dropout variational inference and dropedge variational inference in DR-GST. However, such a weighted loss function will enlarge the impact of incorrect pseudo labels. As a result, we apply the loss correction method to improve the quality of pseudo labels. Both our theoretical analysis and extensive experiments on five benchmark datasets demonstrate the effectiveness of the proposed DR-GST, as well as each well-designed component in DR-GST.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06ca6675-870c-4341-a8fb-4c6970da093eCited by top-tier papers14
- Demystifying Structural Disparity in Graph Neural Networks: Can One Size Fit All?Haitao Mao, Zhikai Chen, Wei Jin, Haoyu Han et al.NeurIPS 2023 · 58 citations
- Deep Insights into Noisy Pseudo Labeling on Graph DataBotao Wang, Jia Li, Yang Liu, Jiashun Cheng et al.NeurIPS 2023 · 24 citations
- Mutually-paced Knowledge Distillation for Cross-lingual Temporal Knowledge Graph ReasoningRuijie Wang, Zheng Li, Jingfeng Yang, Tianyu Cao et al.WWW 2023 · 15 citations
- Empowering Graph Representation Learning with Test-Time Graph TransformationWei Jin, Tong Zhao, Jiayuan Ding, Yozen Liu et al.ICLR 2023 · 10 citations
- Tail-STEAK: Improve Friend Recommendation for Tail Users via Self-Training Enhanced Knowledge DistillationYijun Ma, Chaozhuo Li, Xiao ZhouAAAI 2024 · 10 citations
Builds on8
- DropEdge: Towards Deep Graph Convolutional Networks on Node ClassificationYu Rong, Wenbing Huang, Tingyang Xu, Junzhou HuangICLR 2020 · 1,599 citations
- Multi-Stage Self-Supervised Learning for Graph Convolutional Networks on Graphs with Few Labeled NodesKe Sun, Zhouchen Lin, Zhanxing ZhuAAAI 2020 · 304 citations
- Revisiting Self-Training for Neural Sequence GenerationJunxian He, Jiatao Gu, Jiajun Shen, Marc'Aurelio RanzatoICLR 2020 · 294 citations
- Uncertainty-aware Self-training for Few-shot Text ClassificationSubhabrata Mukherjee, Ahmed Hassan AwadallahNeurIPS 2020 · 182 citations
- Be Confident! Towards Trustworthy Graph Neural Networks via Confidence CalibrationXiao Wang, Hongrui Liu, Chuan Shi, Cheng YangNeurIPS 2021 · 158 citations
Related papers
- Can Pseudo-Label Be More Reliable? A Simple yet Effective Topology-Aware Graph Self-Training MethodGen Liu, Zhongying Zhao, Hui Zhou, Chao Li et al.AAAI 2026
- Defending Graph Convolutional Networks against Dynamic Graph Perturbations via Bayesian Self-SupervisionJun Zhuang, Mohammad Al HasanAAAI 2022 · 48 citations
- Reliable Data Distillation on Graph Convolutional NetworkWentao Zhang, Xupeng Miao, Yingxia Shao, Jiawei Jiang et al.SIGMOD 2020 · 72 citations
- Multi-teacher Self-training for Semi-supervised Node Classification with Noisy LabelsYujing Liu, Zongqian Wu, Zhengyu Lu, Guoqiu Wen et al.ACM MM 2023 · 9 citations
- Adversarial Graph Augmentation to Improve Graph Contrastive LearningSusheel Suresh, Pan Li, Cong Hao, Jennifer NevilleNeurIPS 2021 · 475 citations
