ECBD: Evidence-Centered Benchmark Design for NLP
Yu Lu Liu, Su Lin Blodgett, Jackie C. K. Cheung, Vera Liao, Alexandra Olteanu, Ziang Xiao
摘要
Benchmarking is seen as critical to assessing progress in NLP. However, creating a benchmark involves many design decisions (e.g., which datasets to include, which metrics to use) that often rely on tacit, untested assumptions about what the benchmark is intended to measure or is actually measuring. There is currently no principled way of analyzing these decisions and how they impact the validity of the benchmark's measurements. To address this gap, we draw on evidence-centered design in educational assessments and propose Evidence-Centered Benchmark Design (ECBD), a framework which formalizes the benchmark design process into five modules. ECBD specifies the role each module plays in helping practitioners collect evidence about capabilities of interest. Specifically, each module requires benchmark designers to describe, justify, and support benchmark design choices-e.g., clearly specifying the capabilities the benchmark aims to measure or how evidence about those capabilities is collected from model responses. To demonstrate the use of ECBD, we conduct case studies with three benchmarks: BoolQ, SuperGLUE, and HELM. Our analysis reveals common trends in benchmark design and documentation that could threaten the validity of benchmarks' measurements.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- When AI Benchmarks Plateau: A Systematic Study of Benchmark SaturationMubashara Akhtar, Anka Reuel, Prajna Soni, Sanchit Ahuja 等ICML 2026 · 被引用 22 次
- Can LLMs Reason Abstractly Over Math Word Problems Without CoT? Disentangling Abstract Formulation From Arithmetic ComputationZiling Cheng, Meng Cao, Leila Pishdad, Yanshuai Cao 等EMNLP 2025 · 被引用 6 次
- Learning Task-Agnostic Representations through Multi-Teacher DistillationPhilippe Formont, Maxime Darrin, Banafsheh Karimian, Eric Granger 等NeurIPS 2025 · 被引用 6 次
- Measurement as Bricolage: Examining How Data Scientists Construct Target Variables for Predictive Modeling TasksLuke Guerdan, Devansh Saxena, Stevie Chancellor, Zhiwei Steven Wu 等CSCW 2025 · 被引用 3 次
- Results-Actionability Gap: Understanding How Practitioners Evaluate LLM Products in the WildWillem van der Maden, Malak Sadek, Ziang Xiao, Aske Mottelson 等CHI 2026 · 被引用 2 次
它引用的顶会 Paper9
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- How Robust are Model Rankings : A Leaderboard Customization Approach for Equitable EvaluationSwaroop Mishra, Anjana ArunkumarAAAI 2021 · 被引用 27 次
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 被引用 19 次
- Privacy Implications of Retrieval-Based Language ModelsYangsibo Huang, Samyak Gupta, Zexuan Zhong, Kai Li 等EMNLP 2023 · 被引用 13 次
相关 Paper
- In Benchmarks We Trust ... Or Not?Ine Gevers, Victor De Marez, Jens Van Nooten, Jens Lemmens 等EMNLP 2025 · 被引用 1 次
- E2EDev: Benchmarking Large Language Models in End-to-End Software Development TaskJingyao Liu, Chen Huang, Zhizhao Guan, Wenqiang Lei 等ACL 2026 · 被引用 5 次
- ECBench: Can Multi-modal Foundation Models Understand the Egocentric World? A Holistic Embodied Cognition BenchmarkRonghao Dang, Yuqian Yuan, Wenqi Zhang, Yifei Xin 等CVPR 2025
- Evaluation Examples are not Equally Informative: How should that change NLP Leaderboards?Pedro Rodriguez, Joe Barrow, Alexander Miserlis Hoyle, John P. Lalor 等ACL 2021
- Stereotyping Norwegian Salmon: An Inventory of Pitfalls in Fairness Benchmark DatasetsSu Lin Blodgett, Gilsinia Lopez, Alexandra Olteanu, Robert Sim 等ACL 2021
