First Steps Toward Understanding the Extrapolation of Nonlinear Models to Unseen Domains
Kefan Dong, Tengyu Ma
Abstract
Real-world machine learning applications often involve deploying neural networks to domains that are not seen in the training time. Hence, we need to understand the extrapolation of nonlinear models -- under what conditions on the distributions and function class, models can be guaranteed to extrapolate to new test distributions. The question is very challenging because even two-layer neural networks cannot be guaranteed to extrapolate outside the support of the training distribution without further assumptions on the domain shift. This paper makes some initial steps toward analyzing the extrapolation of nonlinear models for structured domain shift. We primarily consider settings where the marginal distribution of each coordinate of the data (or subset of coordinates) does not shift significantly across the training and test distributions, but the joint distribution may have a much bigger shift. We prove that the family of nonlinear models of the form , where is an arbitrary function on the subset of features , can extrapolate to unseen distributions, if the covariance of the features is well-conditioned. To the best of our knowledge, this is the first result that goes beyond linear models and the bounded density ratio assumption, even though the assumptions on the distribution shift and function class are stylized.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers10
- Compositional Generalization from First PrinciplesThaddäus Wiedemer, Prasanna Mayilvahanan, Matthias Bethge, Wieland BrendelNeurIPS 2023 · 78 citations
- Towards Calibrated Robust Fine-Tuning of Vision-Language ModelsChangdae Oh, Hyesu Lim, Mijoo Kim, Dongyoon Han et al.NeurIPS 2024 · 49 citations
- Provable Compositional Generalization for Object-Centric LearningThaddäus Wiedemer, Jack Brady, Alexander Panfilov, Attila Juhos et al.ICLR 2024 · 40 citations
- Towards Understanding Extrapolation: a Causal LensLingjing Kong, Guangyi Chen, Petar Stojanov, Haoxuan Li et al.NeurIPS 2024 · 7 citations
- Statistical Learning under Heterogenous Distribution ShiftMax Simchowitz, Anurag Ajay, Pulkit Agrawal, Akshay KrishnamurthyICML 2023 · 2 citations
Builds on11
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Theoretical Analysis of Self-Training with Deep Networks on Unlabeled DataColin Wei, Kendrick Shen, Yining Chen, Tengyu MaICLR 2021 · 261 citations
- Extending the WILDS Benchmark for Unsupervised AdaptationShiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao et al.ICLR 2022 · 116 citations
- Connect, Not Collapse: Explaining Contrastive Learning for Unsupervised Domain AdaptationKendrick Shen, Robbie M. Jones, Ananya Kumar, Sang Michael Xie et al.ICML 2022 · 102 citations
- Self-training Avoids Using Spurious Features Under Domain ShiftYining Chen, Colin Wei, Ananya Kumar, Tengyu MaNeurIPS 2020 · 100 citations
Related papers
- How Neural Networks Extrapolate: From Feedforward to Graph Neural NetworksKeyulu Xu, Mozhi Zhang, Jingling Li, Simon Shaolei Du et al.ICLR 2021 · 364 citations
- Out-of-Distribution Generalization via Risk Extrapolation (REx)David Krueger, Ethan Caballero, Jörn-Henrik Jacobsen, Amy Zhang et al.ICML 2021 · 1,163 citations
- Temporal Generalization: A Reality CheckDivyam Madaan, Sumit Chopra, Kyunghyun ChoICLR 2026
- Neural Networks for Learning Counterfactual G-Invariances from Single EnvironmentsS. Chandra Mouli, Bruno RibeiroICLR 2021 · 13 citations
- Distribution Shift Is Key to Learning Invariant PredictionHong Zheng, Fei TengAAAI 2026
