Signal and Noise: A Framework for Reducing Uncertainty in Language Model Evaluation
David Heineman, Valentin Hofmann, Ian Magnusson, Yuling Gu, Noah A. Smith, Hanna Hajishirzi, Kyle Lo, Jesse Dodge
Abstract
Developing large language models is expensive and involves making decisions with small experiments, typically by evaluating on large, multi-task evaluation suites. In this work, we analyze specific properties which make a benchmark more reliable for such decisions, and interventions to design higher-quality evaluation benchmarks. We introduce two key metrics that show differences in current benchmarks: signal, a benchmark's ability to separate better models from worse models, and noise, a benchmark's sensitivity to random variability between training steps. We demonstrate that benchmarks with a better signal-to-noise ratio are more reliable when making decisions at small scale, and those with less noise have lower scaling law prediction error. These results suggest that improving signal or noise will lead to more useful benchmarks, so we introduce three interventions designed to directly affect signal or noise. For example, we propose that switching to a metric that has better signal and noise (e.g., perplexity rather than accuracy) leads to better reliability and improved scaling law error. We also find that filtering noisy subtasks, to improve an aggregate signal-to-noise ratio, leads to more reliable multi-task evaluations. We also find that averaging the output of a model's intermediate checkpoints to reduce noise leads to consistent improvements. We conclude by recommending that those creating new benchmarks, or selecting which existing benchmarks to use, aim for high signal and low noise. We use 30 benchmarks for these experiments, and 375 open-weight language models from 60M to 32B parameters, resulting in a new, publicly available dataset of 900K evaluation benchmark results, totaling 200M instances.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 836db923-4d7a-4f89-b1ff-38bc6a18a01aCited by top-tier papers9
- How Learning Rate Decay Wastes Your Best Data in Curriculum-Based LLM PretrainingKairong Luo, Zhenbo Sun, Haodong Wen, Xinyu Shi et al.ICLR 2026 · 12 citations
- Olmix: A Framework for Data Mixing Throughout LM DevelopmentMayee Chen, Tyler Murray, David Heineman, Matt Jordan et al.ICML 2026 · 9 citations
- When LLMs get significantly worse: A statistical approach to detect model degradationsJonas M. Kübler, Kailash Budhathoki, Matthäus Kleindessner, Xiong Zhou et al.ICLR 2026 · 6 citations
- Train Once, Answer All: Many Pretraining Experiments for the Cost of OneSebastian Bordt, Martin PawelczykICLR 2026 · 6 citations
- Mapping Overlaps in Benchmarks through Perplexity in the WildSiyang Wu, Honglin Bao, Sida Li, Ari Holtzman et al.ICLR 2026 · 5 citations
Builds on19
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 3,037 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
Related papers
- How Reliable is Language Model Micro-Benchmarking?Gregory Yauney, Shahzaib Saqib Warraich, Swabha SwayamdiptaICLR 2026 · 7 citations
- DataDecide: How to Predict Best Pretraining Data with Small ExperimentsIan Magnusson, Nguyen Tai, Ben Bogin, David Heineman et al.ICML 2025
- HE-SNR: Uncovering Latent Logic via Entropy for Guiding Mid-Training on SWE-benchYueyang Wang, Jiawei Fu, Baolong Bi, Xili Wang et al.ICML 2026 · 2 citations
- Revisiting the Scaling Properties of Downstream Metrics in Large Language Model TrainingJakub Krajewski, Amitis Shidani, Dan Busbridge, Sam Wiseman et al.ICLR 2026 · 8 citations
- seqBench: A Tunable Benchmark to Quantify Sequential Reasoning Limits of LLMsMohammad Ramezanali, Mo Vazifeh, Paolo SantiEMNLP 2025
