Path-conditioned training: a principled way to rescale ReLU neural networks
Arthur Lebeurrier, Titouan Vayer, Rémi Gribonval
摘要
Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network parameters. While two properly rescaled weights implement the same function, the training dynamics can be dramatically different. To offer a fresh perspective on exploiting this phenomenon, we build on the recent path-lifting framework, which provides a compact factorization of ReLU networks. We introduce a geometrically motivated criterion to rescale neural network parameters which minimization leads to a conditioning strategy that aligns a kernel in the path-lifting space with a chosen reference. We derive an efficient algorithm to perform this alignment. In the context of random network initialization, we analyze how the architecture and the initialization scale jointly impact the output of the proposed method. Numerical experiments illustrate its potential to speed up training.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins 等ICLR 2021 · 被引用 100 次
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth 等ICML 2021 · 被引用 85 次
- Abide by the law and follow the flow: conservation laws for gradient flowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréNeurIPS 2023 · 被引用 54 次
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen 等NeurIPS 2024 · 被引用 48 次
相关 Paper
- Intrinsic training dynamics of deep neural networksSibylle Marcotte, Gabriel Peyré, Rémi GribonvalICLR 2026 · 被引用 4 次
- Functional vs. parametric equivalence of ReLU networksMary Phuong, Christoph H. LampertICLR 2020 · 被引用 53 次
- A Rescaling-Invariant Lipschitz Bound Based on Path-Metrics for Modern ReLU Network ParameterizationsAntoine Gonon, Nicolas Brisebarre, Elisa Riccietti, Rémi GribonvalICML 2025
- Bounding the Width of Neural Networks via Coupled Initialization A Worst Case AnalysisAlexander Munteanu, Simon Omlor, Zhao Song, David P. WoodruffICML 2022 · 被引用 17 次
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 被引用 35 次
