Gradient Descent as Loss Landscape Navigation: a Normative Framework for Deriving Learning Rules
John J. Vastola, Samuel J. Gershman, Kanaka Rajan
摘要
Learning rules-prescriptions for updating model parameters to improve performance-are typically assumed rather than derived. Why do some learning rules work better than others, and under what assumptions can a given rule be considered optimal? We propose a theoretical framework that casts learning rules as policies for navigating (partially observable) loss landscapes, and identifies optimal rules as solutions to an associated optimal control problem. A range of well-known rules emerge naturally within this framework under different assumptions: gradient descent from short-horizon optimization, momentum from longer-horizon planning, natural gradients from accounting for parameter space geometry, non-gradient rules from partial controllability, and adaptive optimizers like Adam from online Bayesian inference of loss landscape shape. We further show that continual learning strategies like weight resetting can be understood as optimal responses to task uncertainty. By unifying these phenomena under a single objective, our framework clarifies the computational structure of learning and offers a principled foundation for designing adaptive algorithms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper9
- On the Generalization Benefit of Noise in Stochastic Gradient DescentSamuel L. Smith, Erich Elsen, Soham DeICML 2020 · 被引用 122 次
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt 等NeurIPS 2024 · 被引用 100 次
- Noether's Learning Dynamics: Role of Symmetry Breaking in Neural NetworksHidenori Tanaka, Daniel KuninNeurIPS 2021 · 被引用 57 次
- Addressing Loss of Plasticity and Catastrophic Forgetting in Continual LearningMohamed Elsayed, A. Rupam MahmoodICLR 2024 · 被引用 52 次
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 被引用 43 次
相关 Paper
- Understanding Optimization in Deep Learning with Central FlowsJeremy Cohen, Alex Damian, Ameet Talwalkar, J. Zico Kolter 等ICLR 2025
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methodsLaurence AitchisonNeurIPS 2020 · 被引用 23 次
- Noise and Fluctuation of Finite Learning Rate Stochastic Gradient DescentKangqiao Liu, Liu Ziyin, Masahito UedaICML 2021 · 被引用 46 次
- Nested Learning: The Illusion of Deep Learning ArchitecturesAli Behrouz, Meisam Razaviyayn, Peilin Zhong, Vahab MirrokniNeurIPS 2025 · 被引用 96 次
- Reverse engineering learned optimizers reveals known and novel mechanismsNiru Maheswaranathan, David Sussillo, Luke Metz, Ruoxi Sun 等NeurIPS 2021 · 被引用 27 次
