Debiasing Mini-Batch Quadratics for Applications in Deep Learning
Lukas Tatzel, Bálint Mucsányi, Osane Hackel, Philipp Hennig
摘要
Quadratic approximations form a fundamental building block of machine learning methods. E.g., second-order optimizers try to find the Newton step into the minimum of a local quadratic proxy to the objective function; and the second-order approximation of a network's loss function can be used to quantify the uncertainty of its outputs via the Laplace approximation. When computations on the entire training set are intractable-typical for deep learning-the relevant quantities are computed on mini-batches. This, however, distorts and biases the shape of the associated stochastic quadratic approximations in an intricate way with detrimental effects on applications. In this paper, we (i) show that this bias introduces a systematic error, (ii) provide a theoretical explanation for it, (iii) explain its relevance for second-order optimization and uncertainty quantification via the Laplace approximation in deep learning, and (iv) develop and evaluate debiasing strategies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper7
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Being Bayesian, Even Just a Bit, Fixes Overconfidence in ReLU NetworksAgustinus Kristiadi, Matthias Hein, Philipp HennigICML 2020 · 被引用 344 次
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- BackPACK: Packing more into BackpropFelix Dangel, Frederik Kunstner, Philipp HennigICLR 2020 · 被引用 114 次
- Kronecker-Factored Approximate Curvature for Modern Neural Network ArchitecturesRuna Eschenhagen, Alexander Immer, Richard E. Turner, Frank Schneider 等NeurIPS 2023 · 被引用 62 次
相关 Paper
- Exploiting weight-space symmetries for approximating curvatureArtem Artemev, Rui Xia, Benjamin M. Boyd, Youjing Yu 等ICML 2026
- Non-Asymptotic Uncertainty Quantification in High-Dimensional LearningFrederik Hoppe, Claudio Mayrink Verdun, Hannah Laus, Felix Krahmer 等NeurIPS 2024 · 被引用 5 次
- Influence Functions in Deep Learning Are FragileSamyadeep Basu, Phillip Pope, Soheil FeiziICLR 2021 · 被引用 15 次
- Strength of Minibatch Noise in SGDLiu Ziyin, Kangqiao Liu, Takashi Mori, Masahito UedaICLR 2022 · 被引用 44 次
- Exploring Jacobian Inexactness in Second-Order Methods for Variational Inequalities: Lower Bounds, Optimal Algorithms and Quasi-Newton ApproximationsArtem Agafonov, Petr Ostroukhov, Roman Mozhaev, Konstantin Yakovlev 等NeurIPS 2024 · 被引用 6 次
