Learning by Turning: Neural Architecture Aware Optimisation
Yang Liu, Jeremy Bernstein, Markus Meister, Yisong Yue
Abstract
Descent methods for deep networks are notoriously capricious: they require careful tuning of step size, momentum and weight decay, and which method will work best on a new benchmark is a priori unclear. To address this problem, this paper conducts a combined study of neural architecture and optimisation, leading to a new optimiser called Nero: the neuronal rotator. Nero trains reliably without momentum or weight decay, works in situations where Adam and SGD fail, and requires little to no learning rate tuning. Also, Nero's memory footprint is square root that of Adam or LAMB. Nero combines two ideas: (1) projected gradient descent over the space of balanced networks; (2) neuron-specific updates, where the step size sets the angle through which each neuron's hyperplane turns. The paper concludes by discussing how this geometric connection between architecture and optimisation may impact theories of generalisation in deep learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d27bd650-76c5-408b-afd5-0e9ceb1e1a2fCited by top-tier papers10
- Tuning Large Neural Networks via Zero-Shot Hyperparameter TransferGe Yang, Edward J. Hu, Igor Babuschkin, Szymon Sidor et al.NeurIPS 2021 · 208 citations
- Descending through a Crowded Valley - Benchmarking Deep Learning OptimizersRobin M. Schmidt, Frank Schneider, Philipp HennigICML 2021 · 195 citations
- LyaNet: A Lyapunov Framework for Training Neural ODEsIvan Dario Jimenez Rodriguez, Aaron D. Ames, Yisong YueICML 2022 · 78 citations
- Normalization and effective learning rates in reinforcement learningClare Lyle, Zeyu Zheng, Khimya Khetarpal, James Martens et al.NeurIPS 2024 · 69 citations
- Rotational Equilibrium: How Weight Decay Balances Learning Across Neural NetworksAtli Kosson, Bettina Messmer, Martin JaggiICML 2024 · 39 citations
Builds on6
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 1,861 citations
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- On the distance between two neural networks and the stability of learningJeremy Bernstein, Arash Vahdat, Yisong Yue, Ming-Yu LiuNeurIPS 2020 · 77 citations
- Learning compositional functions via multiplicative weight updatesJeremy Bernstein, Jiawei Zhao, Markus Meister, Ming-Yu Liu et al.NeurIPS 2020 · 37 citations
- Characterizing signal propagation to close the performance gap in unnormalized ResNetsAndrew Brock, Soham De, Samuel L. SmithICLR 2021 · 21 citations
Related papers
- Win: Weight-Decay-Integrated Nesterov Acceleration for Adaptive Gradient AlgorithmsPan Zhou, Xingyu Xie, Shuicheng YanICLR 2023
- Adaptive Inertia: Disentangling the Effects of Adaptive Learning Rate and MomentumZeke Xie, Xinrui Wang, Huishuai Zhang, Issei Sato et al.ICML 2022 · 65 citations
- Optimizer Choice Matters For The Emergence of Neural CollapseJim Zhao, Tin Sum Cheng, Wojciech Masarczyk, Aurelien LucchiICLR 2026 · 1 citation
- An Exploration of Non-Euclidean Gradient Descent: Muon and its Many VariantsMichael Crawshaw, Chirag Modi, Mingrui Liu, Robert GowerICML 2026 · 24 citations
- Demystify Hyperparameters for Stochastic Optimization with Transferable RepresentationsJianhui Sun, Mengdi Huai, Kishlay Jha, Aidong ZhangKDD 2022 · 5 citations
