Reliable and Efficient Amortized Model-based Evaluation
Sang T. Truong, Yuheng Tu, Percy Liang, Bo Li, Sanmi Koyejo
Abstract
Comprehensive evaluations of language models (LM) during both development and deployment phases are necessary because these models possess numerous capabilities (e.g., mathematical reasoning, legal support, or medical diagnostic) as well as safety risks (e.g., racial bias, toxicity, or misinformation). The average score across a wide range of benchmarks provides a signal that helps guide the use of these LMs in practice. Currently, holistic evaluations are costly due to the large volume of benchmark questions, making frequent evaluations impractical. A popular attempt to lower the cost is to compute the average score on a subset of the benchmark. This approach, unfortunately, often renders an unreliable measure of LM performance because the average score is often confounded with the difficulty of the questions in the benchmark subset. Item response theory (IRT) was designed to address this challenge, providing a reliable measurement by careful controlling for question difficulty. Unfortunately, question difficulty is expensive to estimate. Facing this challenge, we train a model that predicts question difficulty from its content, enabling a reliable measurement at a fraction of the cost. In addition, we leverage this difficulty predictor to further improve the evaluation efficiency through training a question generator given a difficulty level. This question generator is essential in adaptive testing, where, instead of using a random subset of the benchmark questions, informative questions are adaptively chosen based on the current estimation of LLM performance. Experiments on 22 common natural language benchmarks and 172 LMs show that this approach is more reliable and efficient compared to current common practice. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 84f7b55c-c0c7-4c21-a093-a355cf5ac0ceCited by top-tier papers4
- Lost in Benchmarks? Rethinking Large Language Model Benchmarking with Item Response TheoryHongli Zhou, Hui Huang, Ziqing Zhao, Lvyuan Han et al.AAAI 2026 · 15 citations
- BRIDGE: Predicting Human Task Completion Time From Model PerformanceFengyuan Liu, Jay Gala, Nilaksh, Dzmitry Bahdanau et al.ICML 2026 · 5 citations
- Item Response Scaling Laws: A Measurement Theory Approach for Efficient and Generalizable Neural Scaling EstimationSang Truong, Yuheng Tu, Rylan Schaeffer, Sanmi KoyejoICML 2026
- Benchmarking at the Edge of ComprehensionSamuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb et al.ICML 2026
Builds on3
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor et al.ACL 2021
- Comparing Test Sets with Item Response TheoryClara Vania, Phu Mon Htut, William Huang, Dhara A. Mungra et al.ACL 2021
Related papers
- Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static BenchmarksPeiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng et al.ICML 2026 · 23 citations
- AutoJudger: An Agent-Driven Framework for Efficient Benchmarking of MLLMsXuanwen Ding, Chengjun Pan, Zejun Li, Jiwen Zhang et al.ACL 2026 · 1 citation
- Automated Evaluation of Retrieval-Augmented Language Models with Task-Specific Exam GenerationGauthier Guinet, Behrooz Omidvar-Tehrani, Anoop Deoras, Laurent CallotICML 2024 · 35 citations
- Evaluating Cross-Modal Reasoning Ability and Problem Characteristics with Multimodal Item Response TheoryShunki Uebayashi, Kento Masui, Kyohei Atarashi, Han Bao et al.ICLR 2026 · 1 citation
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba et al.ICML 2025
