In Benchmarks We Trust ... Or Not?
Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens, Andriy Kosar, Ehsan Lotfi, Nikolay Banar, Pieter Fivez, Luna De Bruyne, Walter Daelemans
摘要
Standardized benchmarks are central to evaluating and comparing model performance in Natural Language Processing (NLP). However, Large Language Models (LLMs) have exposed shortcomings in existing benchmarks, and so far there is no clear solution. In this paper, we survey a wide scope of benchmarking issues, and provide an overview of solutions as they are suggested in the literature. We observe that these solutions often tackle a limited number of issues, neglecting other facets. Therefore, we propose concrete checklists to cover all aspects of benchmarking issues, both for benchmark creation and usage. We illustrate the use of our checklists by applying them to three popular NLP benchmarks (i.e., Super-GLUE, WinoGrande, and ARC-AGI). Additionally, we discuss the potential advantages of adding minimal-sized test-suites to benchmarking, which would ensure downstream applicability on real-world use cases.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
- Toward a Perspectivist Turn in Ground Truthing for Predictive ComputingFederico Cabitza, Andrea Campagner, Valerio BasileAAAI 2023 · 被引用 236 次
- Competency Problems: On Finding and Removing Artifacts in Language DataMatt Gardner, William Merrill, Jesse Dodge, Matthew E. Peters 等EMNLP 2021 · 被引用 72 次
- A Systematic Survey and Critical Review on Evaluating Large Language Models: Challenges, Limitations, and RecommendationsMd. Tahmid Rahman Laskar, Sawsan Alqahtani, M. Saiful Bari, Mizanur Rahman 等EMNLP 2024 · 被引用 47 次
- Not All Countries Celebrate Thanksgiving: On the Cultural Dominance in Large Language ModelsWenxuan Wang, Wenxiang Jiao, Jingyuan Huang, Ruyi Dai 等ACL 2024 · 被引用 21 次
相关 Paper
- What's the Meaning of Superhuman Performance in Today's NLU?Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 等ACL 2023 · 被引用 12 次
- Superlim: A Swedish Language Understanding Evaluation BenchmarkAleksandrs Berdicevskis, Gerlof Bouma, Robin Kurtz, Felix Morger 等EMNLP 2023 · 被引用 2 次
- metabench - A Sparse Benchmark of Reasoning and Knowledge in Large Language ModelsAlexander Kipnis, Konstantinos Voudouris, Luca M. Schulze Buschoff, Eric SchulzICLR 2025
- LexGLUE: A Benchmark Dataset for Legal Language Understanding in EnglishIlias Chalkidis, Abhik Jana, Dirk Hartung, Michael J. Bommarito II 等ACL 2022
- ProfBench: Multi-Domain Rubrics requiring Professional Knowledge to Answer and JudgeZhilin Wang, Jaehun Jung, Ximing Lu, Shizhe Diao 等ICLR 2026 · 被引用 20 次
