A Statistical Mechanics Framework for Task-Agnostic Sample Design in Machine Learning
Bhavya Kailkhura, Jayaraman J. Thiagarajan, Qunwei Li, Jize Zhang, Yi Zhou, Timo Bremer
Abstract
In this paper, we present a statistical mechanics framework to understand the effect of sampling properties of training data on the generalization gap of machine learning (ML) algorithms. We connect the generalization gap to the spatial properties of a sample design characterized by the pair correlation function (PCF). In particular, we express generalization gap in terms of the power spectra of the sample design and that of the function to be learned. Using this framework, we show that space-filling sample designs, such as blue noise and Poisson disk sampling, which optimize spectral properties, outperform random designs in terms of the generalization gap and characterize this gain in a closed-form. Our analysis also sheds light on design principles for constructing optimal task-agnostic sample designs that minimize the generalization gap. We corroborate our findings using regression experiments with neural networks on: a) synthetic functions, and b) a complex scientific simulator for inertial confinement fusion (ICF). © 2020 Neural information processing systems foundation. All rights reserved.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9f6da9e4-2c19-4491-a306-fe4be7cc90baRelated papers
- Structure Preserving Neural Networks: A Case Study in the Entropy Closure of the Boltzmann EquationSteffen Schotthöfer, Tianbai Xiao, Martin Frank, Cory D. HauckICML 2022 · 14 citations
- Spectrum Dependent Learning Curves in Kernel Regression and Wide Neural NetworksBlake Bordelon, Abdulkadir Canatar, Cengiz PehlevanICML 2020 · 245 citations
- What Do We Mean by Generalization in Federated Learning?Honglin Yuan, Warren Richard Morningstar, Lin Ning, Karan SinghalICLR 2022 · 98 citations
- Interpretability and Generalization Bounds for Learning Spatial PhysicsAlejandro Queiruga, Theo Gutman-Solo, Shuai JiangICML 2026
- Taxonomizing local versus global structure in neural network loss landscapesYaoqing Yang, Liam Hodgkinson, Ryan Theisen, Joe Zou et al.NeurIPS 2021 · 51 citations
