Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
John J. Vastola, Samuel J. Gershman, Kanaka Rajan
Abstract
Learning rules-prescriptions for updating model parameters to improve performance-are typically assumed rather than derived. Why do some learning rules work better than others, and under what assumptions can a given rule be considered optimal? We propose a theoretical framework that casts learning rules as policies for navigating (partially observable) loss landscapes, and identifies optimal rules as solutions to an associated optimal control problem. A range of well-known rules emerge naturally within this framework under different assumptions: gradient descent from short-horizon optimization, momentum from longer-horizon planning, natural gradients from accounting for parameter space geometry, non-gradient rules from partial controllability, and adaptive optimizers like Adam from online Bayesian inference of loss landscape shape. We further show that continual learning strategies like weight resetting can be understood as optimal responses to task uncertainty. By unifying these phenomena under a single objective, our framework clarifies the computational structure of learning and offers a principled foundation for designing adaptive algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62a1cbdf-bb5b-400e-b3ee-a0146d6443d5Cited by top-tier papers1
Ask how each one uses itBuilds on9
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 122 citations
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt et al.NeurIPS 2024 · 100 citations
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural NetworksHidenori Tanaka, Daniel KuninNeurIPS 2021 · 57 citations
- Addressing Loss of Plasticity and Catastrophic Forgetting in Continual LearningMohamed Elsayed, A. Rupam MahmoodICLR 2024 · 52 citations
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
Related papers
- Understanding Optimization in Deep Learning with Central FlowsJeremy Cohen, Alex Damian, Ameet Talwalkar, J. Zico Kolter et al.ICLR 2025
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methodsLaurence AitchisonNeurIPS 2020 · 23 citations
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient DescentKangqiao Liu, Liu Ziyin, Masahito UedaICML 2021 · 46 citations
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 96 citations
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun et al.NeurIPS 2021 · 27 citations
