Toward Equation of Motion for Deep Neural Networks: Continuous-time Gradient Descent and Discretization Error Analysis
Taiki Miyagawa
Abstract
We derive and solve an ``Equation of Motion'' (EoM) for deep neural networks (DNNs), a differential equation that precisely describes the discrete learning dynamics of DNNs. Differential equations are continuous but have played a prominent role even in the study of discrete optimization (gradient descent (GD) algorithms). However, there still exist gaps between differential equations and the actual learning dynamics of DNNs due to discretization error. In this paper, we start from gradient flow (GF) and derive a counter term that cancels the discretization error between GF and GD. As a result, we obtain EoM, a continuous differential equation that precisely describes the discrete learning dynamics of GD. We also derive discretization error to show to what extent EoM is precise. In addition, we apply EoM to two specific cases: scale- and translation-invariant layers. EoM highlights differences between continuous-time and discrete-time GD, indicating the importance of the counter term for a better description of the discrete learning dynamics of GD. Our experimental results support our theoretical findings.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b94085ad-d032-44b9-8377-4b87f244467aCited by top-tier papers5
- On the Implicit Bias of AdamMatias D. Cattaneo, Jason M. Klusowski, Boris ShigidaICML 2024 · 26 citations
- How does PDE order affect the convergence of PINNs?Changhoon Song, Yesom Park, Myungjoo KangNeurIPS 2024 · 17 citations
- How Memory in Optimization Algorithms Implicitly Modifies the LossMatias D. Cattaneo, Boris ShigidaNeurIPS 2025 · 6 citations
- Heavy-Ball Momentum Method in Continuous Time and Discretization Error AnalysisBochen Lyu, Xiaojing Zhang, Fangyi Zheng, He Wang et al.NeurIPS 2025 · 1 citation
- Improving Explicit Dynamic Gaussian Splatting Optimization via Update MixtureRenjie Ding, Yaonan Wang, Min Liu, Jialin Zhu et al.ICML 2026
Builds on11
- An Exponential Learning Rate Schedule for Deep LearningZhiyuan Li, Sanjeev AroraICLR 2020 · 267 citations
- On the Origin of Implicit Regularization in Stochastic Gradient DescentSamuel L. Smith, Benoit Dherin, David G. T. Barrett, Soham DeICLR 2021 · 235 citations
- Implicit Gradient RegularizationDavid G. T. Barrett, Benoit DherinICLR 2021 · 235 citations
- What Happens after SGD Reaches Zero Loss? --A Mathematical FrameworkZhiyuan Li, Tianhao Wang, Sanjeev AroraICLR 2022 · 121 citations
- Neural Mechanics: Symmetry and Broken Conservation Laws in Deep Learning DynamicsDaniel Kunin, Javier Sagastuy-Breña, Surya Ganguli, Daniel L. K. Yamins et al.ICLR 2021 · 100 citations
Related papers
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 51 citations
- Implicit regularization of deep residual networks towards neural ODEsPierre Marion, Yu-Han Wu, Michael Eli Sander, Gérard BiauICLR 2024 · 24 citations
- Generalization bounds for neural ordinary differential equations and deep residual networksPierre MarionNeurIPS 2023 · 37 citations
- Deep Energy-based Modeling of Discrete-Time PhysicsTakashi Matsubara, Ai Ishikawa, Takaharu YaguchiNeurIPS 2020 · 46 citations
- Learning Efficient and Robust Ordinary Differential Equations via Invertible Neural NetworksWeiming Zhi, Tin Lai, Lionel Ott, Edwin V. Bonilla et al.ICML 2022 · 26 citations
