The Cost of Scaling Down Large Language Models: Reducing Model Size Affects Memory before In-context Learning
Tian Jin, Nolan Clement, Xin Dong, Vaishnavh Nagarajan, Michael Carbin, Jonathan Ragan-Kelley, Gintare Karolina Dziugaite
摘要
Historically, the progress of artificial intelligence (AI) has been mercurial, alternating between surprising breakthroughs and periods of comparative drought. While that will continue to be true to some extent, the incentives and trends are clearer than they have been, and there is even a subfield dedicated to mathematically predicting how much progress to expect from additional investments. These predictions suggest that further performance gains will come from increasing the scale of investment in the current approaches, but that there are sharply diminishing returns. For example, simply increasing the computing budget from 100 million increased the pass rate for AI-generated computer programs from about 65% to about 75%. A billion-dollar version of the model would apparently only reach about 80%, and a trillion-dollar version only 90%. However, the record pass rates already exceed those numbers because users are more inventive in how they apply existing models. This highlights how researcher ingenuity can outperform large-scale investment. Continued investment still has its place. Even shrinking jumps in performance from scaling may continue to justify the rising costs. It may also be that only relatively small improvements in performance "unlock" valuable or risky capabilities that justify policymaker intervention to either enable or avoid them. There are initial signs, though, that these diminishing marginal returns are already dampening the drive for ever-larger models. The growth in compute to train the largest models appears to be slowing. Among publicly available models, users and developers prefer to download intermediate-scale models even when larger and more powerful ones are freely available from the same providers. These trends raise questions about the types of models that will be most impactful and the relative importance of the compute, data, and algorithms that governments might hope to control. They also raise questions about the need for policymakers to intervene to promote or impede that progress and their opportunities to do so.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper1
相关 Paper
- Resource Democratization: Is Compute the Binding Constraint on AI Research?Rebecca Gelles, Veronica Kinoshita, Micah Musser, James DunhamAAAI 2024 · 被引用 4 次
- Algorithmic progress in language modelsAnson Ho, Tamay Besiroglu, Ege Erdil, Zifan Carl Guo 等NeurIPS 2024 · 被引用 51 次
- Language models scale reliably with over-training and on downstream tasksSamir Yitzhak Gadre, Georgios Smyrnis, Vaishaal Shankar, Suchin Gururangan 等ICLR 2025 · 被引用 3 次
- Do Large Language Models Know What They Are Capable Of?Casey O. Barkan, Sidney Black, Oliver SourbutICLR 2026 · 被引用 11 次
- Beyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling LawsNikhil Sardana, Jacob P. Portes, Sasha Doubov, Jonathan FrankleICML 2024 · 被引用 144 次
