How Do Large Language Monkeys Get Their Power (Laws)?
Rylan Schaeffer, Joshua Kazdan, John Hughes, Jordan Juravsky, Sara Price, Aengus Lynch, Erik Jones, Robert Kirk, Azalia Mirhoseini, Sanmi Koyejo
摘要
Recent research across mathematical problem solving, proof assistant programming and multimodal jailbreaking documents a striking finding: when (multimodal) language model tackle a suite of tasks with multiple attempts per task -succeeding if any attempt is correct -then the negative log of the average success rate scales a power law in the number of attempts. In this work, we identify an apparent puzzle: a simple mathematical calculation predicts that on each problem, the failure rate should fall exponentially with the number of attempts. We confirm this prediction empirically, raising a question: from where does aggregate polynomial scaling emerge? We then answer this question by demonstrating per-problem exponential scaling can be made consistent with aggregate polynomial scaling if the distribution of singleattempt success probabilities is heavy tailed such that a small fraction of tasks with extremely low success probabilities collectively warp the aggregate success trend into a power law -even as each problem scales exponentially on its own. We further demonstrate that this distributional perspective explains previously observed deviations from power law scaling, and provides a simple method for forecasting the power law exponent with an order of magnitude lower relative error, or equivalently, ∼2 -4 orders of magnitude less inference compute. Overall, our work contributes to a better understanding of how neural language model performance improves with scaling inference compute and the development of scaling-predictable evaluations of (multimodal) language models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Wider or Deeper? Scaling LLM Inference-Time Compute with Adaptive Branching Tree SearchYuichi Inoue, Kou Misaki, Yuki Imajuku, So Kuroki 等NeurIPS 2025 · 被引用 67 次
- Rethinking Fine-Tuning when Scaling Test-Time Compute: Limiting Confidence Improves Mathematical ReasoningFeng Chen, Allan Raventós, Nan Cheng, Surya Ganguli 等NeurIPS 2025 · 被引用 39 次
- Adaptive Attacks on Trusted Monitors Subvert AI Control ProtocolsMikhail Terekhov, Alexander Panfilov, Daniil Dzenhaliou, Caglar Gulcehre 等ICLR 2026 · 被引用 26 次
- Don’t Pass@k: A Bayesian Framework for Large Language Model EvaluationMohsen Hariri, Amirhossein Samandar, Michael Hinczewski, Vipin ChaudharyICLR 2026 · 被引用 18 次
- Maximum Likelihood Reinforcement LearningFahim Tajwar, Guanning Zeng, Yueer Zhou, Yuda Song 等ICML 2026 · 被引用 18 次
它引用的顶会 Paper42
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Pythia: A Suite for Analyzing Large Language Models Across Training and ScalingStella Biderman, Hailey Schoelkopf, Quentin Gregory Anthony, Herbie Bradley 等ICML 2023 · 被引用 1,822 次
- Scaling Vision Transformers to 22 Billion ParametersMostafa Dehghani, Josip Djolonga, Basil Mustafa, Piotr Padlewski 等ICML 2023 · 被引用 848 次
- Are Emergent Abilities of Large Language Models a Mirage?Rylan Schaeffer, Brando Miranda, Sanmi KoyejoNeurIPS 2023 · 被引用 796 次
- Beyond neural scaling laws: beating power law scaling via data pruningBen Sorscher, Robert Geirhos, Shashank Shekhar, Surya Ganguli 等NeurIPS 2022 · 被引用 720 次
相关 Paper
- A Simple Model of Inference Scaling LawsNoam Itzhak LeviICML 2025
- Learning Shrinks the Hard Tail: Training‑Dependent Inference Scaling in a Solvable Linear ModelNoam Itzhak LeviICLR 2026 · 被引用 7 次
- The Power of Power Law: Asymmetry Enables Compositional ReasoningZixuan Wang, Xingyu Dang, Jason Lee, Kaifeng LyuICML 2026 · 被引用 1 次
- Random Scaling of Emergent CapabilitiesRosie Zhao, Tian Qin, David Alvarez-Melis, Sham Kakade 等ICML 2026 · 被引用 3 次
- Provable Scaling Laws for the Test-Time Compute of Large Language ModelsYanxi Chen, Xuchen Pan, Yaliang Li, Bolin Ding 等NeurIPS 2025 · 被引用 16 次
