Learning compositional functions via multiplicative weight updates
Jeremy Bernstein, Jiawei Zhao, Markus Meister, Ming-Yu Liu, Anima Anandkumar, Yisong Yue
Abstract
Compositionality is a basic structural feature of both biological and artificial neural networks. Learning compositional functions via gradient descent incurs well known problems like vanishing and exploding gradients, making careful learning rate tuning essential for real-world applications. This paper proves that multiplicative weight updates satisfy a descent lemma tailored to compositional functions. Based on this lemma, we derive Madam---a multiplicative version of the Adam optimiser---and show that it can train state of the art neural network architectures without learning rate tuning. We further show that Madam is easily adapted to train natively compressed neural networks by representing their weights in a logarithmic number system. We conclude by drawing connections between multiplicative weight updates and recent findings about synapses in biology.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ed43f0ed-d02f-42a8-9e24-4b3176b92188Cited by top-tier papers5
- Learning by Turning: Neural Architecture Aware OptimisationYang Liu, Jeremy Bernstein, Markus Meister, Yisong YueICML 2021 · 32 citations
- An approach to generate correctly rounded math libraries for new floating point variantsJay P. Lim, Mridul Aanjaneya, John L. Gustafson, Santosh NagarakattePOPL 2021 · 21 citations
- Hyperbolic Aware Minimization: Implicit Bias for SparsityTom Jacobs, Advait Gadhikar, Celia Rubio-Madrigal, Rebekka BurkholzICLR 2026 · 3 citations
- Log-Normal Multiplicative Dynamics for Stable Low-Precision Deep LearningKeigo Nishida, Eren Mehmet KIRAL, Kenichi Bannai, Mohammad Emtiyaz Khan et al.ICML 2026 · 2 citations
- M+Adam: Low-Precision Training via Additive–Multiplicative OptimizationXiaoyuan Liang, Sebastian Loeschcke, Mads Toftrup, Anima AnandkumarICML 2026 · 1 citation
Builds on3
- Large Batch Optimization for Deep Learning: Training BERT in 76 minutesYang You, Jing Li, Sashank J. Reddi, Jonathan Hseu et al.ICLR 2020 · 1,170 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- On the distance between two neural networks and the stability of learningJeremy Bernstein, Arash Vahdat, Yisong Yue, Ming-Yu LiuNeurIPS 2020 · 77 citations
Related papers
- MTAdam: Automatic Balancing of Multiple Training Loss TermsItzik Malkiel, Lior WolfEMNLP 2021 · 12 citations
- Bayesian filtering unifies adaptive and non-adaptive neural network optimization methodsLaurence AitchisonNeurIPS 2020 · 23 citations
- Scalable Optimization in the Modular NormTim Large, Yang Liu, Jacob Huh, Hyojin Bahng et al.NeurIPS 2024 · 70 citations
- Improving Deep Learning Speed and Performance Through Synaptic Neural BalanceAntonios Alexos, Ian Domingo, Pierre BaldiAAAI 2025
- Structured Inverse-Free Natural Gradient Descent: Memory-Efficient & Numerically-Stable KFACWu Lin, Felix Dangel, Runa Eschenhagen, Kirill Neklyudov et al.ICML 2024 · 7 citations
