Prescriptive Scaling Reveals the Evolution of Language Model Capabilities
Hanlin Zhang, Jikai Jin, Vasilis Syrgkanis, Sham Kakade
Abstract
Machine learning model performance improvements tend to arise from competition and application. For deployment, we consider prescriptive scaling laws: given a pre-training compute budget, what downstream accuracy is attainable with contemporary post-training practice, and how stable is that mapping as the field evolves? Using large-scale observational evaluations with 5k existing and 2k newly evaluated model checkpoints spanning 2022-2026 across six benchmarks, we estimate capability boundaries-high conditional quantiles of benchmark scores as a function of log pre-training FLOPs, via smoothed quantile regression with a monotone, saturating sigmoid parameterization. We validate temporal reliability by fitting on earlier model generations and evaluating on later releases: across four of six tasks the out-of-distribution coverage error remains below 2%, while math reasoning exhibits a consistently advancing boundary over time. For instance, at a budget of 10 24 FLOPs the estimated attainable accuracies are 0.83 on IFEval and 0.54 on MATH Lvl 5. We then extend our approach to analyze task-dependent saturation and to probe contamination-related shifts on math reasoning tasks. Finally, we introduce a balanced I-optimal sampling algorithm that recovers near-full-data frontiers using roughly 20% of the parameter-count-weighted evaluation budget (as low as 5% on some tasks) while maintaining comparable calibration. Together, our work releases the Proteus-2k, the latest model performance evaluation dataset, and introduces a practical methodology for translating compute budgets into reliable performance expectations and for monitoring when capability boundaries shift across time. Blog Datasets Code
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fcf5ae49-2519-4e25-bb22-9bb3b744cfeeBuilds on16
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- RL on Incorrect Synthetic Data Scales the Efficiency of LLM Math Reasoning by Eight-FoldAmrith Setlur, Saurabh Garg, Xinyang Geng, Naman Garg et al.NeurIPS 2024 · 143 citations
- Selecting Large Language Model to Fine-tune via Rectified Scaling LawHaowei Lin, Baizhou Huang, Haotian Ye, Qinyu Chen et al.ICML 2024 · 32 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
Related papers
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?Rylan Schaeffer, Hailey Schoelkopf, Brando Miranda, Gabriel Mukobi et al.ICML 2025
- Scaling Behaviors of LLM Reinforcement Learning Post-Training: An Empirical Study in Mathematical ReasoningZelin Tan, Hejia Geng, Xiaohang Yu, Mulei Zhang et al.ACL 2026 · 17 citations
- Broken Neural Scaling LawsEthan Caballero, Kshitij Gupta, Irina Rish, David KruegerICLR 2023 · 15 citations
- Pretraining Scaling Laws for Generative Evaluations of Language ModelsRylan Schaeffer, Noam Levi, Brando Miranda, Sanmi KoyejoICLR 2026 · 5 citations
