Learning compositional functions via multiplicative weight updates
Jeremy Bernstein, Jiawei Zhao, Markus Meister, Ming-Yu Liu, Anima Anandkumar, Yisong Yue
摘要
Compositionality is a basic structural feature of both biological and artificial neural networks. Learning compositional functions via gradient descent incurs well known problems like vanishing and exploding gradients, making careful learning rate tuning essential for real-world applications. This paper proves that multiplicative weight updates satisfy a descent lemma tailored to compositional functions. Based on this lemma, we derive Madam---a multiplicative version of the Adam optimiser---and show that it can train state of the art neural network architectures without learning rate tuning. We further show that Madam is easily adapted to train natively compressed neural networks by representing their weights in a logarithmic number system. We conclude by drawing connections between multiplicative weight updates and recent findings about synapses in biology.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning by Turning: Neural Architecture Aware OptimisationYang Liu, Jeremy Bernstein, Markus Meister, Yisong YueICML 2021 · 被引用 32 次
- An approach to generate correctly rounded math libraries for new floating point variantsJay P. Lim, Mridul Aanjaneya, John L. Gustafson, Santosh NagarakattePOPL 2021 · 被引用 21 次
- Hyperbolic Aware Minimization: Implicit Bias for SparsityTom Jacobs, Advait Gadhikar, Celia Rubio-Madrigal, Rebekka BurkholzICLR 2026 · 被引用 3 次
- Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep LearningKeigo Nishida, Eren Mehmet KIRAL, Kenichi Bannai, Mohammad Emtiyaz Khan 等ICML 2026 · 被引用 2 次
- M+Adam: Low-Precision Training via Additive–Multiplicative OptimizationXiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup, Anima AnandkumarICML 2026 · 被引用 1 次
它引用的顶会 Paper3
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu 等ICLR 2020 · 被引用 1,170 次
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 被引用 598 次
- On the distance between two neural networks and the stability of learningJeremy Bernstein, Arash Vahdat, Yisong Yue, Ming-Yu LiuNeurIPS 2020 · 被引用 77 次
相关 Paper
- MTAdam: Automatic Balancing of Multiple Training Loss TermsItzik Malkiel, Lior WolfEMNLP 2021 · 被引用 12 次
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methodsLaurence AitchisonNeurIPS 2020 · 被引用 23 次
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng 等NeurIPS 2024 · 被引用 70 次
- Improving Deep Learning Speed and Performance Through Synaptic Neural BalanceAntonios Alexos, Ian Domingo, Pierre BaldiAAAI 2025
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov 等ICML 2024 · 被引用 7 次
