Path-conditioned training: a principled way to rescale ReLU neural networks
Arthur Lebeurrier, Titouan Vayer, Rémi Gribonval
Abstract
Despite recent algorithmic advances, we still lack principled ways to leverage the well-documented rescaling symmetries in ReLU neural network parameters. While two properly rescaled weights implement the same function, the training dynamics can be dramatically different. To offer a fresh perspective on exploiting this phenomenon, we build on the recent path-lifting framework, which provides a compact factorization of ReLU networks. We introduce a geometrically motivated criterion to rescale neural network parameters which minimization leads to a conditioning strategy that aligns a kernel in the path-lifting space with a chosen reference. We derive an efficient algorithm to perform this alignment. In the context of random network initialization, we analyze how the architecture and the initialization scale jointly impact the output of the proposed method. Numerical experiments illustrate its potential to speed up training.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 38e10c79-77f6-415e-b613-5be21c6db4c4Cited by top-tier papers1
Ask how each one uses itBuilds on13
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins et al.ICLR 2021 · 100 citations
- On the Implicit Bias of Initialization Shape: Beyond Infinitesimal Mirror DescentShahar Azulay, Edward Moroshko, Mor Shpigel Nacson, Blake E. Woodworth et al.ICML 2021 · 85 citations
- Abide by the law and follow the flow: conservation laws for gradient flowsSibylle Marcotte, Rémi Gribonval, Gabriel PeyréNeurIPS 2023 · 54 citations
- Get rich quick: exact solutions reveal how unbalanced initializations promote rapid feature learningDaniel Kunin, Allan Raventós, Clémentine C. J. Dominé, Feng Chen et al.NeurIPS 2024 · 48 citations
Related papers
- Intrinsic training dynamics of deep neural networksSibylle Marcotte, Gabriel Peyré, Rémi GribonvalICLR 2026 · 4 citations
- Functional vs. parametric equivalence of ReLU networksMary Phuong, Christoph H. LampertICLR 2020 · 53 citations
- A Rescaling-Invariant Lipschitz Bound Based on Path-Metrics for Modern ReLU Network ParameterizationsAntoine Gonon, Nicolas Brisebarre, Elisa Riccietti, Rémi GribonvalICML 2025
- Bounding the Width of Neural Networks via Coupled Initialization A Worst Case AnalysisAlexander Munteanu, Simon Omlor, Zhao Song, David P. WoodruffICML 2022 · 17 citations
- Hidden Symmetries of ReLU NetworksJ. Elisenda Grigsby, Kathryn Lindsey, David RolnickICML 2023 · 35 citations
