Learning by Turning: Neural Architecture Aware Optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, Yisong Yue
摘要
Descent methods for deep networks are notoriously capricious: they require careful tuning of step size, momentum and weight decay, and which method will work best on a new benchmark is a priori unclear. To address this problem, this paper conducts a combined study of neural architecture and optimisation, leading to a new optimiser called Nero: the neuronal rotator. Nero trains reliably without momentum or weight decay, works in situations where Adam and SGD fail, and requires little to no learning rate tuning. Also, Nero's memory footprint is square root that of Adam or LAMB. Nero combines two ideas: (1) projected gradient descent over the space of balanced networks; (2) neuron-specific updates, where the step size sets the angle through which each neuron's hyperplane turns. The paper concludes by discussing how this geometric connection between architecture and optimisation may impact theories of generalisation in deep learning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor 等NeurIPS 2021 · 被引用 208 次
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 被引用 195 次
- LyaNet: A Lyapunov Framework for Training Neural ODEsIvan Dario Jimenez Rodriguez, Aaron D. Ames, Yisong YueICML 2022 · 被引用 78 次
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens 等NeurIPS 2024 · 被引用 69 次
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 被引用 39 次
它引用的顶会 Paper6
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- On the distance between two neural networks and the stability of learningJeremy Bernstein, Arash Vahdat, Yisong Yue, Ming-Yu LiuNeurIPS 2020 · 被引用 77 次
- Learning compositional functions via multiplicative weight updatesJeremy Bernstein, Jiawei Zhao, Markus Meister, Ming-Yu Liu 等NeurIPS 2020 · 被引用 37 次
- Characterizing signal propagation to close the performance gap in unnormalized ResNetsAndrew Brock, Soham De, Samuel L. SmithICLR 2021 · 被引用 21 次
相关 Paper
- Win: Weight-Decay-Integrated Nesterov Acceleration for Adaptive Gradient AlgorithmsPan Zhou, Xingyu Xie, Shuicheng YanICLR 2023
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato 等ICML 2022 · 被引用 65 次
- Optimizer Choice Matters For The Emergence of Neural CollapseJim Zhao, Tin Sum Cheng, Wojciech Masarczyk, Aurelien LucchiICLR 2026 · 被引用 1 次
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many VariantsMichael Crawshaw, Chirag Modi, Mingrui Liu, Robert GowerICML 2026 · 被引用 24 次
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 被引用 5 次
