Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language Models
Frederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt, Alberto Bietti
Abstract
Adam has been shown to outperform gradient descent on large language models by a larger margin than on other tasks, but it is unclear why. We show that a key factor in this performance gap is the heavy-tailed class imbalance found in language tasks. When trained with gradient descent, the loss of infrequent words decreases more slowly than the loss of frequent ones. This leads to a slow decrease on the average loss as most samples come from infrequent words. On the other hand, Adam and sign-based methods are less sensitive to this problem. To establish that this behavior is caused by class imbalance, we show empirically that it can be reproduced across architectures and data types, on language transformers, vision CNNs, and linear models. On a linear model with cross-entropy loss, we show that class imbalance leads to imbalanced, correlated gradients and Hessians that have been hypothesized to benefit Adam. We also prove that, in continuous time, gradient descent converges slowly on low-frequency classes while sign descent does not.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers33
- Small Batch Size Training for Language Models: When Vanilla SGD Works, and Why Gradient Accumulation is WastefulMartin Marek, Sanae Lotfi, Aditya Somasundaram, Andrew Gordon Wilson et al.NeurIPS 2025 · 46 citations
- In Search of Adam's Secret SauceAntonio Orvieto, Robert GowerNeurIPS 2025 · 43 citations
- Muon Outperforms Adam in Tail-End Associative Memory LearningShuche Wang, Fengzhuo Zhang, Jiaxiang Li, Cunxiao Du et al.ICLR 2026 · 40 citations
- Adam with model exponential moving average is effective for nonconvex optimizationKwangjun Ahn, Ashok CutkoskyNeurIPS 2024 · 36 citations
- Understanding and Minimising Outlier Features in Transformer TrainingBobby He, Lorenzo Noci, Daniele Paliotta, Imanol Schlag et al.NeurIPS 2024 · 27 citations
Builds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Long-tail learning via logit adjustmentAditya Krishna Menon, Sadeep Jayasumana, Ankit Singh Rawat, Himanshu Jain et al.ICLR 2021 · 937 citations
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real et al.NeurIPS 2023 · 734 citations
- Why Gradient Clipping Accelerates Training: A Theoretical Justification for AdaptivityJingzhao Zhang, Tianxing He, Suvrit Sra, Ali JadbabaieICLR 2020 · 598 citations
- Why are Adaptive Methods Good for Attention Models?Jingzhao Zhang, Sai Praneeth Karimireddy, Andreas Veit, Seungyeon Kim et al.NeurIPS 2020 · 397 citations
Related papers
- Noise Is Not the Main Factor Behind the Gap Between Sgd and Adam on Transformers, But Sign Descent Might BeFrederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, Mark SchmidtICLR 2023 · 5 citations
- Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's LawFrederik Kunstner, Francis BachNeurIPS 2025 · 21 citations
- Why Transformers Need Adam: A Hessian PerspectiveYushun Zhang, Congliang Chen, Tian Ding, Ziniu Li et al.NeurIPS 2024 · 149 citations
- Understanding Adam Requires Better Rotation Dependent AssumptionsTianyue H. Zhang, Lucas Maes, Alan Milligan, Alexia Jolicoeur-Martineau et al.NeurIPS 2025 · 11 citations
- Frequency-aware SGD for Efficient Embedding Learning with Provable BenefitsYan Li, Dhruv Choudhary, Xiaohan Wei, Baichuan Yuan et al.ICLR 2022 · 7 citations
