Why Propagate Alone? Parallel Use of Labels and Features on Graphs
Yangkun Wang, Jiarui Jin, Weinan Zhang, Yongyi Yang, Jiuhai Chen, Quan Gan, Yong Yu, Zheng Zhang, Zengfeng Huang, David Wipf
Abstract
Graph neural networks (GNNs) and label propagation represent two interrelated modeling strategies designed to exploit graph structure in tasks such as node property prediction. The former is typically based on stacked message-passing layers that share neighborhood information to transform node features into predictive embeddings. In contrast, the latter involves spreading label information to unlabeled nodes via a parameter-free diffusion process, but operates independently of the node features. Given then that the material difference is merely whether features or labels are smoothed across the graph, it is natural to consider combinations of the two for improving performance. In this regard, it has recently been proposed to use a randomly-selected portion of the training labels as GNN inputs, concatenated with the original node features for making predictions on the remaining labels. This so-called label trick accommodates the parallel use of features and labels, and is foundational to many of the top-ranking submissions on the Open Graph Benchmark (OGB) leaderboard. And yet despite its wide-spread adoption, thus far there has been little attempt to carefully unpack exactly what statistical properties the label trick introduces into the training pipeline, intended or otherwise. To this end, we prove that under certain simplifying assumptions, the stochastic label trick can be reduced to an interpretable, deterministic training objective composed of two factors. The first is a data-fitting term that naturally resolves potential label leakage issues, while the second serves as a regularization factor conditioned on graph structure that adapts to graph size and connectivity. Later, we leverage this perspective to motivate a broader range of label trick use cases, and provide experiments to verify the efficacy of these extensions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on5
- Open Graph Benchmark: Datasets for Machine Learning on GraphsWeihua Hu, Matthias Fey, Marinka Zitnik, Yuxiao Dong et al.NeurIPS 2020 · 3,935 citations
- Combining Label Propagation and Simple Models out-performs Graph Neural NetworksQian Huang, Horace He, Abhay Singh, Ser-Nam Lim et al.ICLR 2021 · 322 citations
- Training Graph Neural Networks with 1000 LayersGuohao Li, Matthias Müller, Bernard Ghanem, Vladlen KoltunICML 2021 · 294 citations
- Graph Neural Networks Inspired by Classical Iterative AlgorithmsYongyi Yang, Tang Liu, Yangkun Wang, Jinjing Zhou et al.ICML 2021 · 94 citations
- Boost then Convolve: Gradient Boosting Meets Graph Neural NetworksSergei Ivanov, Liudmila ProkhorenkovaICLR 2021 · 21 citations
Related papers
- GLASS: GNN with Labeling Tricks for Subgraph Representation LearningXiyuan Wang, Muhan ZhangICLR 2022 · 38 citations
- Labeling Trick: A Theory of Using Graph Neural Networks for Multi-Node Representation LearningMuhan Zhang, Pan Li, Yinglong Xia, Kai Wang et al.NeurIPS 2021 · 255 citations
- Neo-GNNs: Neighborhood Overlap-aware Graph Neural Networks for Link PredictionSeongjun Yun, Seoyoon Kim, Junhyun Lee, Jaewoo Kang et al.NeurIPS 2021 · 183 citations
- Divide and Denoise: Empowering Simple Models for Robust Semi-Supervised Node Classification against Label NoiseKaize Ding, Xiaoxiao Ma, Yixin Liu, Shirui PanKDD 2024 · 8 citations
- A Unified Lottery Ticket Hypothesis for Graph Neural NetworksTianlong Chen, Yongduo Sui, Xuxi Chen, Aston Zhang et al.ICML 2021 · 208 citations
