The Value of Out-of-Distribution Data
Ashwin De Silva, Rahul Ramesh, Carey E. Priebe, Pratik Chaudhari, Joshua T. Vogelstein
Abstract
Generalization error always improves with more in-distribution data. However, it is an open question what happens as we add out-of-distribution (OOD) data. Intuitively, if the OOD data is quite different, it seems more data would harm generalization error, though if the OOD data are sufficiently similar, much empirical evidence suggests that OOD data can actually improve generalization error. We show a counter-intuitive phenomenon: the generalization error of a task can be a non-monotonic function of the amount of OOD data. Specifically, we show that generalization error can improve with small amounts of OOD data, and then get worse with larger amounts compared to no OOD data. In other words, there is value in training on small amounts of OOD data. We analytically demonstrate these results via Fisher's Linear Discriminant on synthetic datasets, and empirically demonstrate them via deep networks on computer vision benchmarks such as MNIST, CIFAR-10, CINIC-10, PACS and DomainNet. In the idealistic setting where we know which samples are OOD, we show that these non-monotonic trends can be exploited using an appropriately weighted objective of the target and OOD empirical risk. While its practical utility is limited, this does suggest that if we can detect OOD samples, then there may be ways to benefit from them. When we do not know which samples are OOD, we show how a number of go-to strategies such as data-augmentation, hyperparameter optimization and pre-training are not enough to ensure that the target generalization error does not deteriorate with the number of OOD samples in the dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Prospective Learning: Learning for a Dynamic FutureAshwin De Silva, Rahul Ramesh, Rubing Yang, Siyu Yu et al.NeurIPS 2024 · 5 citations
- Ensemble Distribution Distillation via Flow MatchingJonggeon Park, Giung Nam, Hyunsu Kim, Jongmin Yoon et al.ICML 2025
Builds on14
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Energy-based Out-of-distribution DetectionWeitang Liu, Xiaoyun Wang, John D. Owens, Yixuan LiNeurIPS 2020 · 2,213 citations
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma et al.ICLR 2022 · 911 citations
- Out-of-Distribution Detection with Deep Nearest NeighborsYiyou Sun, Yifei Ming, Xiaojin Zhu, Yixuan LiICML 2022 · 789 citations
- Exploring the Limits of Out-of-Distribution DetectionStanislav Fort, Jie Ren, Balaji LakshminarayananNeurIPS 2021 · 443 citations
Related papers
- In or Out? Fixing ImageNet Out-of-Distribution Detection EvaluationJulian Bitterwolf, Maximilian Müller, Matthias HeinICML 2023 · 154 citations
- In-N-Out: Pre-Training and Self-Training using Auxiliary Information for Out-of-Distribution RobustnessSang Michael Xie, Ananya Kumar, Robbie Jones, Fereshte Khani et al.ICLR 2021 · 69 citations
- Don't forget the nullspace! Nullspace occupancy as a mechanism for out of distribution failureDaksh Idnani, Vivek Madan, Naman Goyal, David J. Schwab et al.ICLR 2023
- Scaling laws for learning with real and surrogate dataAyush Jain, Andrea Montanari, Eren SasogluNeurIPS 2024 · 30 citations
- Igeood: An Information Geometry Approach to Out-of-Distribution DetectionEduardo Dadalto Câmara Gomes, Florence Alberge, Pierre Duhamel, Pablo PiantanidaICLR 2022 · 32 citations
