On the distance between two neural networks and the stability of learning
Jeremy Bernstein, Arash Vahdat, Yisong Yue, Ming-Yu Liu
Abstract
This paper relates parameter distance to gradient breakdown for a broad class of nonlinear compositional functions. The analysis leads to a new distance function called deep relative trust and a descent lemma for neural networks. Since the resulting learning rule seems to require little to no learning rate tuning, it may unlock a simpler workflow for training deeper and more complex neural networks. The Python code used in this paper is here: https://github.com/jxbz/fromage . But what does ∆G mean? In what sense should it be small? We see that this is another area that could benefit from an appropriate notion of distance on neural networks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers24
- AdaBelief Optimizer: Adapting Stepsizes by the Belief in Observed GradientsJuntang Zhuang, Tommy Tang, Yifan Ding, Sekhar Tatikonda et al.NeurIPS 2020 · 697 citations
- High-Performance Large-Scale Image Recognition Without NormalizationAndy Brock, Soham De, Samuel L. Smith, Karen SimonyanICML 2021 · 613 citations
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- DoG is SGD's Best Friend: A Parameter-Free Dynamic Step Size ScheduleMaor Ivgi, Oliver Hinder, Yair CarmonICML 2023 · 98 citations
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng et al.NeurIPS 2024 · 70 citations
Builds on1
Related papers
- Learning compositional functions via multiplicative weight updatesJeremy Bernstein, Jiawei Zhao, Markus Meister, Ming-Yu Liu et al.NeurIPS 2020 · 37 citations
- How DNNs break the Curse of Dimensionality: Compositionality and Symmetry LearningArthur Jacot, Seok Hoan Choi, Yuxiao WenICLR 2025
- DoWG Unleashed: An Efficient Universal Parameter-Free Gradient Descent MethodAhmed Khaled, Konstantin Mishchenko, Chi JinNeurIPS 2023 · 49 citations
- Toward Equation of Motion for Deep Neural Networks: Continuous-time Gradient Descent and Discretization Error AnalysisTaiki MiyagawaNeurIPS 2022 · 14 citations
- Depth Separation with Multilayer Mean-Field NetworksYunwei Ren, Mo Zhou, Rong GeICLR 2023
