When does preconditioning help or hurt generalization?
Shun-ichi Amari, Jimmy Ba, Roger Baker Grosse, Xuechen Li, Atsushi Nitanda, Taiji Suzuki, Denny Wu, Ji Xu
Abstract
While second order optimizers such as natural gradient descent (NGD) often speed up optimization, their effect on generalization has been called into question. This work presents a more nuanced view on how the implicit bias of first- and second-order methods affects the comparison of generalization properties. We provide an exact asymptotic bias-variance decomposition of the generalization error of overparameterized ridgeless regression under a general class of preconditioner , and consider the inverse population Fisher information matrix (used in NGD) as a particular example. We determine the optimal for both the bias and variance, and find that the relative generalization performance of different optimizers depends on the label noise and the "shape" of the signal (true parameters): when the labels are noisy, the model is misspecified, or the signal is misaligned with the features, NGD can achieve lower risk; conversely, GD generalizes better than NGD under clean labels, a well-specified model, or aligned signal. Based on this analysis, we discuss several approaches to manage the bias-variance tradeoff, and the potential benefit of interpolating between GD and NGD. We then extend our analysis to regression in the reproducing kernel Hilbert space and demonstrate that preconditioned GD can decrease the population risk faster than GD. Lastly, we empirically compare the generalization error of first- and second-order optimizers in neural network experiments, and observe robust trends matching our theoretical analysis.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb9bc019-7969-43c4-bb79-1c96e0deb8aaCited by top-tier papers16
- On the Optimal Weighted Regularization in Overparameterized Linear RegressionDenny Wu, Ji XuNeurIPS 2020 · 151 citations
- An Unconstrained Layer-Peeled Perspective on Neural CollapseWenlong Ji, Yiping Lu, Yiliang Zhang, Zhun Deng et al.ICLR 2022 · 101 citations
- A Unified Framework for U-Net Design and AnalysisChristopher Williams, Fabian Falck, George Deligiannidis, Chris C. Holmes et al.NeurIPS 2023 · 79 citations
- The Power of Preconditioning in Overparameterized Low-Rank Matrix SensingXingyu Xu, Yandi Shen, Yuejie Chi, Cong MaICML 2023 · 51 citations
- Multi-scale Feature Learning Dynamics: Insights for Double DescentMohammad Pezeshki, Amartya Mitra, Yoshua Bengio, Guillaume LajoieICML 2022 · 33 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Deep Double Descent: Where Bigger Models and More Data HurtPreetum Nakkiran, Gal Kaplun, Yamini Bansal, Tristan Yang et al.ICLR 2020 · 1,108 citations
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- Generalisation error in learning with random features and the hidden manifold modelFederica Gerace, Bruno Loureiro, Florent Krzakala, Marc Mézard et al.ICML 2020 · 184 citations
- Implicit Regularization in Deep Learning May Not Be Explainable by NormsNoam Razin, Nadav CohenNeurIPS 2020 · 178 citations
Related papers
- On Optimal Interpolation in Linear RegressionEduard Oravkin, Patrick RebeschiniNeurIPS 2021 · 6 citations
- Implicit Bias of Spectal Descent and Muon on Multiclass Separable DataChen Fan, Mark Schmidt, Christos ThrampoulidisNeurIPS 2025
- The Benefits of Implicit Regularization from SGD in Least Squares ProblemsDifan Zou, Jingfeng Wu, Vladimir Braverman, Quanquan Gu et al.NeurIPS 2021 · 41 citations
- Rich Information is Affordable: A Systematic Performance Analysis of Second-order Optimization Using K-FACYuichiro Ueno, Kazuki Osawa, Yohei Tsuji, Akira Naruse et al.KDD 2020 · 9 citations
- Flat Minima in Linear Estimation and an Extended Gauss Markov TheoremSimon N. SegertICLR 2024 · 1 citation
