Not-Just-Scaling Laws: Towards a Better Understanding of the Downstream Impact of Language Model Design Decisions
Emmy Liu, Amanda Bertsch, Lintang Sutawika, Lindia Tjuatja, Patrick Fernandes, Lara Marinov, Michael Chen, Shreya Singhal, Carolin Lawrence, Aditi Raghunathan, Kiril Gashteovski, Graham Neubig
Abstract
Improvements in language model capabilities are often attributed to increasing model size or training data, but in some cases smaller models trained on curated data or with different architectural decisions can outperform larger ones trained on more tokens. What accounts for this? To quantify the impact of these design choices, we meta-analyze 92 open-source pretrained models across a wide array of scales, including state-of-the-art open-weights models as well as less performant models and those with less conventional design decisions. We find that by incorporating features besides model size and number of training tokens, we can achieve a relative 3-28% increase in ability to predict downstream performance compared with using scale alone. Analysis of model design decisions reveal insights into data composition, such as the trade-off between language and code tasks at 15-25% code, as well as the negative association of web data with truthfulness. Broadly, our framework lays a foundation for more systematic investigation of how model development choices shape final capabilities. 1 Use of AI Assistants Claude 3.5 Sonnet and GPT-o3-mini-high were used to help revise and shorten several parts of this submission as well as to edit for clarity. The first draft was entirely human-written.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5671bcb9-9155-4019-adc2-b80b4bda4fd0Cited by top-tier papers3
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 18 citations
- EvoLM: In Search of Lost Training Dynamics for Language Model ReasoningZhenting Qi, Fan Nie, Alexandre Alahi, James Y. Zou et al.NeurIPS 2025 · 3 citations
- Overtrained Language Models Are Harder to Fine-TuneJacob Mitchell Springer, Sachin Goyal, Kaiyue Wen, Tanishq Kumar et al.ICML 2025
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- On Layer Normalization in the Transformer ArchitectureRuibin Xiong, Yunchang Yang, Di He, Kai Zheng et al.ICML 2020 · 1,388 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- An Information Theoretic Perspective on Agentic System DesignShizhe He, Avanika Narayan, Ishan S. Khare, Scott W. Linderman et al.ICLR 2026 · 6 citations
- Training Trajectories of Language Models Across ScalesMengzhou Xia, Mikel Artetxe, Chunting Zhou, Xi Victoria Lin et al.ACL 2023 · 12 citations
- From Parameters to Performance: A Data-Driven Study on LLM Structure and DevelopmentSuqing Wang, Zuchao Li, Luohe Shi, Bo Du et al.EMNLP 2025
- How much pretraining data do language models need to learn syntax?Laura Pérez-Mayos, Miguel Ballesteros, Leo WannerEMNLP 2021 · 31 citations
- Look Before You Leap: Universal Emergent Mechanism for Retrieval in Language ModelsAlexandre Variengien, Eric WinsorICLR 2025
