Deep Neural Networks Tend To Extrapolate Predictably
Katie Kang, Amrith Setlur, Claire J. Tomlin, Sergey Levine
Abstract
Conventional wisdom suggests that neural network predictions tend to be unpredictable and overconfident when faced with out-of-distribution (OOD) inputs. Our work reassesses this assumption for neural networks with high-dimensional inputs. Rather than extrapolating in arbitrary ways, we observe that neural network predictions often tend towards a constant value as input data becomes increasingly OOD. Moreover, we find that this value often closely approximates the optimal constant solution (OCS), i.e., the prediction that minimizes the average loss over the training data without observing the input. We present results showing this phenomenon across 8 datasets with different distributional shifts (including CIFAR10-C and ImageNet-R, S), different loss functions (cross entropy, MSE, and Gaussian NLL), and different architectures (CNNs and transformers). Furthermore, we present an explanation for this behavior, which we first validate empirically and then study theoretically in a simplified setting involving deep homogeneous networks with ReLU activations. Finally, we show how one can leverage our insights in practice to enable risk-sensitive decision-making in the presence of OOD inputs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Saddle-to-Saddle Dynamics Explains A Simplicity Bias Across Neural Network ArchitecturesYedi Zhang, Andrew M. Saxe, Peter E. LathamICLR 2026 · 15 citations
- Just One Layer Norm Guarantees Stable ExtrapolationJuliusz Ziomek, George Whittle, Michael A. OsborneNeurIPS 2025 · 4 citations
- Quantifying Distributional Invariance in Causal Subgraph for IRM-Free Graph GeneralizationYang Qiu, Yixiong Zou, Jun Wang, Wei Liu et al.NeurIPS 2025 · 3 citations
- Heads collapse, features stay: Why Replay needs big buffersGiulia Lanzillotta, Damiano Meier, Thomas HofmannICLR 2026 · 3 citations
- Fully Heteroscedastic Count Regression with Deep Double Poisson NetworksSpencer Young, Porter Jenkins, Longchao Da, Jeffrey Dotson et al.ICML 2025
Builds on16
- The Many Faces of Robustness: A Critical Analysis of Out-of-Distribution GeneralizationDan Hendrycks, Steven Basart, Norman Mu, Saurav Kadavath et al.ICCV 2021 · 2,294 citations
- WILDS: A Benchmark of in-the-Wild Distribution ShiftsPang Wei Koh, Shiori Sagawa, Henrik Marklund, Sang Michael Xie et al.ICML 2021 · 1,773 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
- CSI: Novelty Detection via Contrastive Learning on Distributionally Shifted InstancesJihoon Tack, Sangwoo Mo, Jongheon Jeong, Jinwoo ShinNeurIPS 2020 · 755 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
Related papers
- Mitigating Neural Network Overconfidence with Logit NormalizationHongxin Wei, Renchunzi Xie, Hao Cheng, Lei Feng et al.ICML 2022 · 386 citations
- Towards neural networks that provably know when they don't knowAlexander Meinke, Matthias HeinICLR 2020 · 151 citations
- Entropy Maximization and Meta Classification for Out-of-Distribution Detection in Semantic SegmentationRobin Chan, Matthias Rottmann, Hanno GottschalkICCV 2021 · 200 citations
- The Risks of Invariant Risk MinimizationElan Rosenfeld, Pradeep Kumar Ravikumar, Andrej RisteskiICLR 2021 · 356 citations
- Robustness via Cross-Domain EnsemblesTeresa Yeo, Oguzhan Fatih Kar, Amir ZamirICCV 2021 · 30 citations
