Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent Generalization
Neha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer, Jascha Sohl-Dickstein
Abstract
Machine learning is predicated on the concept of generalization: a model achieving low error on a sufficiently large training set should also perform well on novel samples from the same distribution. We show that both data whitening and second order optimization can harm or entirely prevent generalization. In general, model training harnesses information contained in the sample-sample second moment matrix of a dataset. For a general class of models, namely models with a fully connected first layer, we prove that the information contained in this matrix is the only information which can be used to generalize. Models trained using whitened data, or with certain second order optimization schemes, have less access to this information, resulting in reduced or nonexistent generalization ability. We experimentally verify these predictions for several architectures, and further demonstrate that generalization continues to be harmed even when theoretical requirements are relaxed. However, we also show experimentally that regularized second order optimization can provide a practical tradeoff, where training is accelerated but less information is lost, and generalization can in some circumstances even improve.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ea61dc8c-c968-4833-95e7-4fc3920ad1d6Cited by top-tier papers8
- Amortized Proximal OptimizationJuhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. GrosseNeurIPS 2022 · 15 citations
- Are ID Embeddings Necessary? Whitening Pre-trained Text Embeddings for Effective Sequential RecommendationLingzi Zhang, Xin Zhou, Zhiwei Zeng, Zhiqi ShenICDE 2024 · 9 citations
- Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor GeneralizationDavide Buffelli, Jamie McGowan, Wangkun Xu, Alexandru Cioba et al.NeurIPS 2024 · 6 citations
- Newton Losses: Using Curvature Information for Learning with Differentiable AlgorithmsFelix Petersen, Christian Borgelt, Tobias Sutter, Hilde Kuehne et al.NeurIPS 2024 · 3 citations
- On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning RegimeShuai Jiang, Eric C. Cyr, Ben S. Southworth, Alexey VoroninICLR 2026 · 1 citation
Builds on4
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam et al.NeurIPS 2020 · 245 citations
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 167 citations
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 127 citations
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li et al.ICLR 2021 · 11 citations
Related papers
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan et al.ICML 2020 · 125 citations
- Neglected Hessian component explains mysteries in sharpness regularizationYann N. Dauphin, Atish Agarwala, Hossein MobahiNeurIPS 2024 · 16 citations
- Why adversarial training can hurt robust accuracyJacob Clarysse, Julia Hörrmann, Fanny YangICLR 2023 · 6 citations
- Reducing Per-Sample Harm in Stochastic OptimizationApostolos AvranasICML 2026
- Debiasing Mini-Batch Quadratics for Applications in Deep LearningLukas Tatzel, Bálint Mucsányi, Osane Hackel, Philipp HennigICLR 2025
