Whitening and Second Order Optimization Both Make Information in the Dataset Unusable During Training, and Can Reduce or Prevent Generalization
Neha S. Wadia, Daniel Duckworth, Samuel S. Schoenholz, Ethan Dyer, Jascha Sohl-Dickstein
摘要
Machine learning is predicated on the concept of generalization: a model achieving low error on a sufficiently large training set should also perform well on novel samples from the same distribution. We show that both data whitening and second order optimization can harm or entirely prevent generalization. In general, model training harnesses information contained in the sample-sample second moment matrix of a dataset. For a general class of models, namely models with a fully connected first layer, we prove that the information contained in this matrix is the only information which can be used to generalize. Models trained using whitened data, or with certain second order optimization schemes, have less access to this information, resulting in reduced or nonexistent generalization ability. We experimentally verify these predictions for several architectures, and further demonstrate that generalization continues to be harmed even when theoretical requirements are relaxed. However, we also show experimentally that regularized second order optimization can provide a practical tradeoff, where training is accelerated but less information is lost, and generalization can in some circumstances even improve.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Amortized Proximal OptimizationJuhan Bae, Paul Vicol, Jeff Z. HaoChen, Roger B. GrosseNeurIPS 2022 · 被引用 15 次
- Are ID Embeddings Necessary? Whitening Pre-trained Text Embeddings for Effective Sequential RecommendationLingzi Zhang, Xin Zhou, Zhiwei Zeng, Zhiqi ShenICDE 2024 · 被引用 9 次
- Exact, Tractable Gauss-Newton Optimization in Deep Reversible Architectures Reveal Poor GeneralizationDavide Buffelli, Jamie McGowan, Wangkun Xu, Alexandru Cioba 等NeurIPS 2024 · 被引用 6 次
- Newton Losses: Using Curvature Information for Learning with Differentiable AlgorithmsFelix Petersen, Christian Borgelt, Tobias Sutter, Hilde Kuehne 等NeurIPS 2024 · 被引用 3 次
- On the Convergence Behavior of Preconditioned Gradient Descent Toward the Rich Learning RegimeShuai Jiang, Eric C. Cyr, Ben S. Southworth, Alexey VoroninICLR 2026 · 被引用 1 次
它引用的顶会 Paper4
- Finite Versus Infinite Neural Networks: an Empirical StudyJaehoon Lee, Samuel S. Schoenholz, Jeffrey Pennington, Ben Adlam 等NeurIPS 2020 · 被引用 245 次
- Dynamics of Deep Neural Networks and Neural Tangent HierarchyJiaoyang Huang, Horng-Tzer YauICML 2020 · 被引用 167 次
- Asymptotics of Wide Networks from Feynman DiagramsEthan Dyer, Guy Gur-AriICLR 2020 · 被引用 127 次
- When does preconditioning help or hurt generalization?Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li 等ICLR 2021 · 被引用 11 次
相关 Paper
- On the Noisy Gradient Descent that Generalizes as SGDJingfeng Wu, Wenqing Hu, Haoyi Xiong, Jun Huan 等ICML 2020 · 被引用 125 次
- Neglected Hessian component explains mysteries in sharpness regularizationYann N. Dauphin, Atish Agarwala, Hossein MobahiNeurIPS 2024 · 被引用 16 次
- Why adversarial training can hurt robust accuracyJacob Clarysse, Julia Hörrmann, Fanny YangICLR 2023 · 被引用 6 次
- Reducing Per-Sample Harm in Stochastic OptimizationApostolos AvranasICML 2026
- Debiasing Mini-Batch Quadratics for Applications in Deep LearningLukas Tatzel, Bálint Mucsányi, Osane Hackel, Philipp HennigICLR 2025
