Understanding Optimization in Deep Learning with Central Flows
Jeremy Cohen, Alex Damian, Ameet Talwalkar, J. Zico Kolter, Jason D. Lee
摘要
Optimization in deep learning remains poorly understood. A key difficulty is that optimizers exhibit complex oscillatory dynamics, referred to as "edge of stability," which cannot be captured by traditional optimization theory. In this paper, we show that the path taken by an oscillatory optimizer can often be captured by a central flow: a differential equation which directly models the time-averaged (i.e. smoothed) optimization trajectory. We empirically show that these central flows can predict long-term optimization trajectories for generic neural networks with a high degree of numerical accuracy. By interpreting these flows, we are able to understand how gradient descent makes progress even as the loss sometimes goes up; how adaptive optimizers ``adapt'' to the local loss landscape; and how adaptive optimizers implicitly seek out regions of weight space where they can take larger steps. These insights (and others) are not apparent from the optimizers' update rules, but are revealed by the central flows. Therefore, we believe that central flows constitute a promising tool for reasoning about optimization in deep learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper23
- Training Dynamics Impact Post-Training Quantization RobustnessAlbert Catalan-Tatjer, Niccolò Ajroldi, Jonas GeipingICLR 2026 · 被引用 13 次
- Large Stepsizes Accelerate Gradient Descent for Regularized Logistic RegressionJingfeng Wu, Pierre Marion, Peter L. BartlettNeurIPS 2025 · 被引用 12 次
- Neural Thermodynamics: Entropic Forces in Deep and Universal Representation LearningLiu Ziyin, Yizhou Xu, Isaac L. ChuangNeurIPS 2025 · 被引用 11 次
- Adaptive Preconditioners Trigger Loss Spikes in AdamZhiwei Bai, Zhangchen Zhou, Jiajie Zhao, Xiaolong Li 等ICML 2026 · 被引用 9 次
- Through the River: Understanding the Benefit of Schedule-Free Methods for Language Model TrainingMinhak Song, Beomhan Baek, Kwangjun Ahn, Chulhee YunNeurIPS 2025 · 被引用 9 次
相关 Paper
- On the Convergence Direction of Gradient DescentShuo Chen, Xiaolong Li, Jiaying Peng, Yao ZhaoICLR 2026
- Gradient Descent Monotonically Decreases the Sharpness of Gradient Flow Solutions in Scalar Networks and BeyondItai Kreisler, Mor Shpigel Nacson, Daniel Soudry, Yair CarmonICML 2023 · 被引用 19 次
- Gradient Descent on Neural Networks Typically Occurs at the Edge of StabilityJeremy Cohen, Simran Kaur, Yuanzhi Li, J. Zico Kolter 等ICLR 2021 · 被引用 22 次
- Flatland: The Adventures of Gradient Descent with Large Step SizesLeonardo Galli, Curtis Fox, Wiebke Bartolomaeus, Mark Schmidt 等ICML 2026
- Continuous vs. Discrete Optimization of Deep Neural NetworksOmer Elkabetz, Nadav CohenNeurIPS 2021 · 被引用 51 次
