In Benchmarks We Trust ... Or Not?
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
Abstract
Standardized benchmarks are central to evaluating and comparing model performance in Natural Language Processing (NLP). However, Large Language Models (LLMs) have exposed shortcomings in existing benchmarks, and so far there is no clear solution. In this paper, we survey a wide scope of benchmarking issues, and provide an overview of solutions as they are suggested in the literature. We observe that these solutions often tackle a limited number of issues, neglecting other facets. Therefore, we propose concrete checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage. We illustrate the use of our checklists by applying them to three popular NLP benchmarks (i.e., Super-GLUE, WinoGrande, and ARC-AGI). Additionally, we discuss the potential advantages of adding minimal-sized test-suites to benchmarking, which would ensure downstream applicability on real-world use cases.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 27ea9c11-c1b3-42b4-b28a-fff6f85ce54dBuilds on9
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 236 citations
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters et al.EMNLP 2021 · 72 citations
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman et al.EMNLP 2024 · 47 citations
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai et al.ACL 2024 · 21 citations
Related papers
- What's the Meaning of Superhuman Performance in Today's NLU?Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic et al.ACL 2023 · 12 citations
- Superlim: A Swedish Language Understanding Evaluation BenchmarkAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger et al.EMNLP 2023 · 2 citations
- metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language ModelsAlexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric SchulzICLR 2025
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II et al.ACL 2022
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and JudgeZhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao et al.ICLR 2026 · 20 citations
