Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?
Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor, Robin Jia, Jordan L. Boyd-Graber
Abstract
Leaderboards are widely used in NLP and push the field forward. While leaderboards are a straightforward ranking of NLP models, this simplicity can mask nuances in evaluation items (examples) and subjects (NLP models). Rather than replace leaderboards, we advocate a re-imagining so that they better highlight if and where progress is made. Building on educational testing, we create a Bayesian leaderboard model where latent subject skill and latent item difficulty predict correct responses. Using this model, we analyze the ranking reliability of leaderboards. Afterwards, we show the model can guide what to annotate, identify annotation errors, detect overfitting, and identify informative examples. We conclude with recommendations for future benchmark tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a7b88d41-a154-4afc-8a0a-234af549641eCited by top-tier papers32
- Understanding Dataset Difficulty with V-Usable InformationKawin Ethayarajh, Yejin Choi, Swabha SwayamdiptaICML 2022 · 337 citations
- tinyBenchmarks: evaluating LLMs with fewer examplesFelipe Maia Polo, Lucas Weber, Leshem Choshen, Yuekai Sun et al.ICML 2024 · 212 citations
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva et al.NeurIPS 2024 · 93 citations
- Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-TuningQin Liu, Rui Zheng, Bao Rong, Jingyi Liu et al.ACL 2022 · 35 citations
- How Benchmark Prediction from Fewer Data Misses the MarkGuanhua Zhang, Florian E. Dorner, Moritz HardtNeurIPS 2025 · 26 citations
Builds on8
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- Identifying Mislabeled Data using the Area Under the Margin RankingGeoff Pleiss, Tianyi Zhang, Ethan R. Elenberg, Kilian Q. WeinbergerNeurIPS 2020 · 398 citations
- Beyond Accuracy: Behavioral Testing of NLP Models with CheckListMarco Túlio Ribeiro, Tongshuang Wu, Carlos Guestrin, Sameer SinghACL 2020 · 51 citations
- What do Models Learn from Question Answering Datasets?Priyanka Sen, Amir SaffariEMNLP 2020 · 40 citations
Related papers
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed et al.ACL 2024 · 13 citations
- How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable EvaluationSwaroop Mishra, Anjana ArunkumarAAAI 2021 · 27 citations
- AI Cartography: Mapping the Latent Landscape of AI Benchmark EcosystemsMichael Hardy, Anka Reuel, Lijin Zhang, Jodi Casabianca et al.ICML 2026
- La Leaderboard: A Large Language Model Leaderboard for Spanish Varieties and Languages of Spain and Latin AmericaMaría Grandury, Javier Aula-Blasco, Júlia Falcão, Clémentine Fourrier et al.ACL 2025
- Dynaboard: An Evaluation-As-A-Service Platform for Holistic Next-Generation BenchmarkingZhiyi Ma, Kawin Ethayarajh, Tristan Thrush, Somya Jain et al.NeurIPS 2021 · 76 citations
