Algorithmic progress in language models
Anson Ho, Tamay Besiroglu, Ege Erdil, Zifan Carl Guo, David Owen, Robi Rahman, David Atkinson, Neil Thompson, Jaime Sevilla
Abstract
We investigate the rate at which algorithms for pre-training language models have improved since the advent of deep learning. Using a dataset of over 200 language model evaluations on Wikitext and Penn Treebank spanning 2012-2023, we find that the compute required to reach a set performance threshold has halved approximately every 8 months, with a 95% confidence interval of around 5 to 14 months, substantially faster than hardware gains per Moore's Law. We estimate augmented scaling laws, which enable us to quantify algorithmic progress and determine the relative contributions of scaling models versus innovations in training algorithms. Despite the rapid pace of algorithmic progress and the development of new architectures such as the transformer, our analysis reveals that the increase in compute made an even larger contribution to overall performance improvements over this time period. Though limited by noisy benchmark data, our analysis quantifies the rapid progress in language modeling, shedding light on the relative contributions from compute and algorithms.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Measuring AI Ability to Complete Long Software TasksThomas Kwa, Ben West, Joel Becker, Amy Deng et al.NeurIPS 2025 · 160 citations
- Fine-grained Uncertainty Decomposition in Large Language Models: A Spectral ApproachNassim Walha, Sebastian G. Gruber, Thomas Decker, Yinchong Yang et al.AAAI 2026 · 2 citations
- Beyond Tokens: Enhancing RTL Quality Estimation via Structural Graph LearningYi Liu, Hongji Zhang, Yiwen Wang, Dimitrios Tsaras et al.ICML 2026 · 1 citation
- DICE: Data Influence Cascade in Decentralized LearningTongtian Zhu, Wenhao Li, Can Wang, Fengxiang HeICLR 2025
Builds on5
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli et al.NeurIPS 2022 · 720 citations
- Scaling Data-Constrained Language ModelsNiklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao et al.NeurIPS 2023 · 475 citations
- Neuro-Symbolic Language Modeling with Automaton-augmented RetrievalUri Alon, Frank F. Xu, Junxian He, Sudipta Sengupta et al.ICML 2022 · 79 citations
Related papers
- Pre-training under infinite computeKonwoo Kim, Suhas Kotha, Percy Liang, Tatsunori HashimotoICLR 2026 · 25 citations
- Cramming: Training a Language Model on a single GPU in one dayJonas Geiping, Tom GoldsteinICML 2023 · 115 citations
- On the origin of neural scaling laws: from random graphs to natural languageMaissam Barkeshli, Alberto Alfarano, Andrey GromovICML 2026
- Parallel Scaling Law for Language ModelsMouxiang Chen, Binyuan Hui, Zeyu Cui, Jiaxi Yang et al.NeurIPS 2025 · 33 citations
- Scaling Properties of Speech Language ModelsSantiago Cuervo, Ricard MarxerEMNLP 2024 · 5 citations
