The Geometry of Neural Nets' Parameter Spaces Under Reparametrization
Agustinus Kristiadi, Felix Dangel, Philipp Hennig
摘要
Model reparametrization, which follows the change-of-variable rule of calculus, is a popular way to improve the training of neural nets. But it can also be problematic since it can induce inconsistencies in, e.g., Hessian-based flatness measures, optimization trajectories, and modes of probability densities. This complicates downstream analyses: e.g. one cannot definitively relate flatness with generalization since arbitrary reparametrization changes their relationship. In this work, we study the invariance of neural nets under reparametrization from the perspective of Riemannian geometry. From this point of view, invariance is an inherent property of any neural net if one explicitly represents the metric and uses the correct associated transformation rules. This is important since although the metric is always present, it is often implicitly assumed as identity, and thus dropped from the notation, then lost under reparametrization. We discuss implications for measuring the flatness of minima, optimization, and for probability-density maximization. Finally, we explore some interesting directions where invariance is useful.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Reparameterization invariance in approximate Bayesian inferenceHrittik Roy, Marco Miani, Carl Henrik Ek, Philipp Hennig 等NeurIPS 2024 · 被引用 20 次
- Riemannian Laplace approximations for Bayesian neural networksFederico Bergamin, Pablo Moreno-Muñoz, Søren Hauberg, Georgios ArvanitidisNeurIPS 2023 · 被引用 18 次
- A Tale of Two Symmetries: Exploring the Loss Landscape of Equivariant ModelsYuqing Xie, Tess E. SmidtNeurIPS 2025 · 被引用 9 次
- Hide & Seek: Transformer Symmetries Obscure Sharpness & Riemannian Geometry Finds ItMarvin F. da Silva, Felix Dangel, Sageev OoreICML 2025
- Sharper Convergence Rates for Nonconvex Optimisation via Reduction MappingsEvan Markou, Thalaiyasingam Ajanthan, Stephen GouldNeurIPS 2025
它引用的顶会 Paper17
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Fantastic Generalization Measures and Where to Find ThemYiding Jiang, Behnam Neyshabur, Hossein Mobahi, Dilip Krishnan 等ICLR 2020 · 被引用 705 次
- Laplace Redux - Effortless Bayesian Deep LearningErik A. Daxberger, Agustinus Kristiadi, Alexander Immer, Runa Eschenhagen 等NeurIPS 2021 · 被引用 508 次
- Tensor Programs IV: Feature Learning in Infinite-Width Neural NetworksGreg Yang, Edward J. HuICML 2021 · 被引用 242 次
- Scalable Marginal Likelihood Estimation for Model Selection in Deep LearningAlexander Immer, Matthias Bauer, Vincent Fortuin, Gunnar Rätsch 等ICML 2021 · 被引用 130 次
相关 Paper
- A Reparametrization-Invariant Sharpness Measure Based on Information GeometryCheongjae Jang, Sungyoon Lee, Frank C. Park, Yung-Kyun NohNeurIPS 2022 · 被引用 19 次
- Relative Flatness and GeneralizationHenning Petzka, Michael Kamp, Linara Adilova, Cristian Sminchisescu 等NeurIPS 2021 · 被引用 114 次
- Deep Networks on Toroids: Removing Symmetries Reveals the Structure of Flat Regions in the Landscape GeometryFabrizio Pittorino, Antonio Ferraro, Gabriele Perugini, Christoph Feinauer 等ICML 2022 · 被引用 30 次
- In What Ways Are Deep Neural Networks Invariant and How Should We Measure This?Henry Kvinge, Tegan Emerson, Grayson Jorgenson, Scott Vasquez 等NeurIPS 2022 · 被引用 15 次
- Normalized Flat Minima: Exploring Scale Invariant Definition of Flat Minima for Neural Networks Using PAC-Bayesian AnalysisYusuke Tsuzuku, Issei Sato, Masashi SugiyamaICML 2020 · 被引用 91 次
