Scaling Laws for Gradient Descent and Sign Descent for Linear Bigram Models under Zipf's Law
Frederik Kunstner, Francis Bach
摘要
Recent works have highlighted optimization difficulties faced by gradient descent in training the first and last layers of transformer-based language models, which are overcome by optimizers such as Adam. These works suggest that the difficulty is linked to the heavy-tailed distribution of words in text data, where the frequency of the th most frequent word is proportional to , following Zipf's law. To better understand the impact of the data distribution on training performance, we study a linear bigram model for next-token prediction when the tokens follow a power law parameterized by the exponent . We derive optimization scaling laws for deterministic gradient descent and sign descent as a proxy for Adam as a function of the exponent . Existing theoretical investigations in scaling laws assume that the eigenvalues of the data decay as a power law with exponent . This assumption effectively makes the problem finite dimensional'' as most of the loss comes from a few of the largest eigencomponents. In comparison, we show that the problem is more difficult when the data have heavier tails. The case $\alpha = 1$ as found in text data is worst-case'' for gradient descent, in that the number of iterations required to reach a small relative error scales almost linearly with dimension. While the performance of sign descent also depends on the dimension, for Zipf-distributed data the number of iterations scales only with the square-root of the dimension, leading to a large improvement for large vocabularies.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Scaling Laws and Spectra of Shallow Neural Networks in the Feature Learning RegimeLeonardo Defilippis, Yizhou Xu, Julius Girardin, Vittorio Erba 等ICLR 2026 · 被引用 20 次
- Theory of Scaling Laws for In-Context Regression: Depth, Width, Context and TimeBlake Bordelon, Mary I. Letey, Cengiz PehlevanICLR 2026 · 被引用 14 次
- Muon in Associative Memory Learning: Training Dynamics and Scaling LawsKaifei Wang, Binghui Li, Han Zhong, Pinyan Lu 等ICML 2026 · 被引用 7 次
- Fast Catch-Up, Late Switching: Optimal Batch Size Scheduling via Functional Scaling LawsJinbo Wang, Binghui Li, Zhanpeng Zhou, Mingze Wang 等ICLR 2026 · 被引用 6 次
- Single-Head Attention in High Dimensions: A Theory of Generalization, Weights Spectra, and Scaling LawsFabrizio Boncoraglio, Vittorio Erba, Emanuele Troiani, Yizhou Xu 等ICML 2026 · 被引用 5 次
它引用的顶会 Paper21
- Symbolic Discovery of Optimization AlgorithmsXiangning Chen, Chen Liang, Da Huang, Esteban Real 等NeurIPS 2023 · 被引用 734 次
- A Constructive Prediction of the Generalization Error Across ScalesJonathan S. Rosenfeld, Amir Rosenfeld, Yonatan Belinkov, Nir ShavitICLR 2020 · 被引用 265 次
- Understanding the Difficulty of Training TransformersLiyuan Liu, Xiaodong Liu, Jianfeng Gao, Weizhu Chen 等EMNLP 2020 · 被引用 158 次
- Scaling Laws with Vocabulary: Larger Models Deserve Larger VocabulariesChaofan Tao, Qian Liu, Longxu Dou, Niklas Muennighoff 等NeurIPS 2024 · 被引用 135 次
- On the SDEs and Scaling Rules for Adaptive Gradient AlgorithmsSadhika Malladi, Kaifeng Lyu, Abhishek Panigrahi, Sanjeev AroraNeurIPS 2022 · 被引用 125 次
相关 Paper
- Heavy-Tailed Class Imbalance and Why Adam Outperforms Gradient Descent on Language ModelsFrederik Kunstner, Alan Milligan, Robin Yadav, Mark Schmidt 等NeurIPS 2024 · 被引用 100 次
- Noise Is Not the Main Factor Behind the Gap Between Sgd and Adam on Transformers, But Sign Descent Might BeFrederik Kunstner, Jacques Chen, Jonathan Wilder Lavington, Mark SchmidtICLR 2023 · 被引用 5 次
- Universal One-third Time Scaling in Learning Peaked DistributionsYizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff GoreICML 2026 · 被引用 6 次
- Zipfian WhiteningSho Yokoi, Han Bao, Hiroto Kurita, Hidetoshi ShimodairaNeurIPS 2024 · 被引用 3 次
- Deconstructing What Makes a Good Optimizer for Autoregressive Language ModelsRosie Zhao, Depen Morwani, David Brandfonbrener, Nikhil Vyas 等ICLR 2025
