Continuous vs. Discrete Optimization of Deep Neural Networks
Omer Elkabetz, Nadav Cohen
Abstract
Existing analyses of optimization in deep learning are either continuous, focusing on (variants of) gradient flow, or discrete, directly treating (variants of) gradient descent. Gradient flow is amenable to theoretical analysis, but is stylized and disregards computational efficiency. The extent to which it represents gradient descent is an open question in the theory of deep learning. The current paper studies this question. Viewing gradient descent as an approximate numerical solution to the initial value problem of gradient flow, we find that the degree of approximation depends on the curvature around the gradient flow trajectory. We then show that over deep neural networks with homogeneous activations, gradient flow trajectories enjoy favorable curvature, suggesting they are well approximated by gradient descent. This finding allows us to translate an analysis of gradient flow over deep linear neural networks into a guarantee that gradient descent efficiently converges to global minimum almost surely under random initialization. Experiments suggest that over simple deep neural networks, gradient descent with conventional step size is indeed close to gradient flow. We hypothesize that the theory of gradient flows will unravel mysteries behind deep learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers20
- Implicit Regularization in Hierarchical Tensor Factorization and Deep Convolutional Neural NetworksNoam Razin, Asaf Maman, Nadav CohenICML 2022 · 34 citations
- Perturbation Analysis of Neural CollapseTom Tirer, Haoxiang Huang, Jonathan Niles-WeedICML 2023 · 32 citations
- DiEP: Adaptive Mixture-of-Experts Compression through Differentiable Expert PruningSikai Bai, Haoxi Li, Jie Zhang, Zicong Hong et al.NeurIPS 2025 · 27 citations
- Vanishing Gradients in Reinforcement Finetuning of Language ModelsNoam Razin, Hattie Zhou, Omid Saremi, Vimal Thilak et al.ICLR 2024 · 27 citations
- On the Effective Number of Linear Regions in Shallow Univariate ReLU Networks: Convergence Guarantees and Implicit BiasItay Safran, Gal Vardi, Jason D. LeeNeurIPS 2022 · 26 citations
Builds on13
- Gradient Descent Maximizes the Margin of Homogeneous Neural NetworksKaifeng Lyu, Jian LiICLR 2020 · 402 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- Directional convergence and alignment in deep learningZiwei Ji, Matus TelgarskyNeurIPS 2020 · 226 citations
- The Break-Even Point on Optimization Trajectories of Deep Neural NetworksStanislaw Jastrzebski, Maciej Szymczak, Stanislav Fort, Devansh Arpit et al.ICLR 2020 · 198 citations
Related papers
- Implicit Bias in Leaky ReLU Networks Trained on High-Dimensional DataSpencer Frei, Gal Vardi, Peter L. Bartlett, Nathan Srebro et al.ICLR 2023 · 5 citations
- Non-Singularity of the Gradient Descent Map for Neural Networks with Piecewise Analytic ActivationsAlexandru Craciun, Debarghya GhoshdastidarNeurIPS 2025 · 1 citation
- A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat MinimaZeke Xie, Issei Sato, Masashi SugiyamaICLR 2021 · 165 citations
- Understanding Optimization in Deep Learning with Central FlowsJeremy Cohen, Alex Damian, Ameet Talwalkar, J. Zico Kolter et al.ICLR 2025
- On the Explicit Role of Initialization on the Convergence and Implicit Bias of Overparametrized Linear NetworksHancheng Min, Salma Tarmoun, René Vidal, Enrique MalladaICML 2021 · 53 citations
