How Neural Networks Extrapolate: From Feedforward to Graph Neural Networks
Keyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du, Ken-ichi Kawarabayashi, Stefanie Jegelka
Abstract
We study how neural networks trained by gradient descent extrapolate, i.e., what they learn outside the support of the training distribution. Previous works report mixed empirical results when extrapolating with neural networks: while feedforward neural networks, a.k.a. multilayer perceptrons (MLPs), do not extrapolate well in certain simple tasks, Graph Neural Networks (GNNs) -structured networks with MLP modules -have shown some success in more complex tasks. Working towards a theoretical explanation, we identify conditions under which MLPs and GNNs extrapolate well. First, we quantify the observation that ReLU MLPs quickly converge to linear functions along any direction from the origin, which implies that ReLU MLPs do not extrapolate most nonlinear functions. But, they can provably learn a linear target function when the training distribution is sufficiently "diverse". Second, in connection to analyzing the successes and limitations of GNNs, these results suggest a hypothesis for which we provide theoretical and empirical evidence: the success of GNNs in extrapolating algorithmic tasks to new data (e.g., larger graphs or edge weights) relies on encoding task-specific non-linearities in the architecture or features. Our theoretical analysis builds on a connection of over-parameterized networks to the neural tangent kernel. Empirically, our theory holds across different training settings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7aa526da-cb29-46f2-87a8-bb7e912fe586Cited by top-tier papers102
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 986 citations
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
- Handling Distribution Shifts on Graphs: An Invariance PerspectiveQitian Wu, Hengrui Zhang, Junchi Yan, David WipfICLR 2022 · 261 citations
- Learning Causally Invariant Representations for Out-of-Distribution Generalization on GraphsYongqiang Chen, Yonggang Zhang, Yatao Bian, Han Yang et al.NeurIPS 2022 · 246 citations
- GraphNorm: A Principled Approach to Accelerating Graph Neural Network TrainingTianle Cai, Shengjie Luo, Keyulu Xu, Di He et al.ICML 2021 · 224 citations
Builds on18
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Domain Generalization with MixStyleKaiyang Zhou, Yongxin Yang, Yu Qiao, Tao XiangICLR 2021 · 986 citations
- Principal Neighbourhood Aggregation for Graph NetsGabriele Corso, Luca Cavalleri, Dominique Beaini, Pietro Liò et al.NeurIPS 2020 · 914 citations
Related papers
- A Non-Parametric Regression Viewpoint : Generalization of Overparametrized Deep RELU Network Under Noisy ObservationsNamjoon Suh, Hyunouk Ko, Xiaoming HuoICLR 2022 · 15 citations
- Graph Neural Networks are Inherently Good Generalizers: Insights by Bridging GNNs and MLPsChenxiao Yang, Qitian Wu, Jiahua Wang, Junchi YanICLR 2023 · 16 citations
- Zero-One Laws of Graph Neural NetworksSam Adam-Day, Theodor-Mihai Iliant, Ismail Ilkan CeylanNeurIPS 2023 · 11 citations
- Why Do Deep Residual Networks Generalize Better than Deep Feedforward Networks? - A Neural Tangent Kernel PerspectiveKaixuan Huang, Yuqing Wang, Molei Tao, Tuo ZhaoNeurIPS 2020 · 107 citations
- Learning to Execute Graph Algorithms Exactly with Graph Neural NetworksMuhammad Fetrat Qharabagh, Artur Back de Luca, George Giapitzakis, Kimon FountoulakisICML 2026
