Examining the robustness of LLM evaluation to the distributional assumptions of benchmarks
Charlotte Siska, Katerina Marazopoulou, Melissa Ailem, James Bono
摘要
Benchmarks have emerged as the central approach for evaluating Large Language Models (LLMs). The research community often relies on a model's average performance across the test prompts of a benchmark to evaluate the model's performance. This is consistent with the assumption that the test prompts within a benchmark represent a random sample from a real-world distribution of interest. We note that this is generally not the case; instead, we hold that the distribution of interest varies according to the specific use case. We find that (1) the correlation in model performance across test prompts is non-random, (2) accounting for correlations across test prompts can change model rankings on major benchmarks, (3) explanatory factors for these correlations include semantic similarity and common LLM failure points.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- Towards Acyclic Preference Evaluation of Language Models via Multiple EvaluatorsZhengyu Hu, Jieyu Zhang, Zhihan Xiong, Alexander Ratner 等AAAI 2026 · 被引用 14 次
- AgentNoiseBench: Benchmarking Robustness of Tool-Using LLM Agents Under Noisy ConditionRuipeng Wang, Yuxin Chen, Yukai Wang, Chang Wu 等ICML 2026 · 被引用 12 次
- Investigating Value-Reasoning Reliability in Small Large Language ModelsXia Du, Shuhan Sun, Pengyuan Liu, Dong YuEMNLP 2025 · 被引用 2 次
- Robust AI Evaluation through Maximal LotteriesHadi Khalaf, Serena Wang, Daniel Halpern, Itai Shapira 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- Quantifying Language Models' Sensitivity to Spurious Features in Prompt Design or: How I learned to start worrying about prompt formattingMelanie Sclar, Yejin Choi, Yulia Tsvetkov, Alane SuhrICLR 2024 · 被引用 682 次
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal 等ACL 2020 · 被引用 602 次
- GLM-130B: An Open Bilingual Pre-trained ModelAohan Zeng, Xiao Liu, Zhengxiao Du, Zihan Wang 等ICLR 2023 · 被引用 295 次
相关 Paper
- Efficient multi-prompt evaluation of LLMsFelipe Maia Polo, Ronald Xu, Lucas Weber, Mírian Silva 等NeurIPS 2024 · 被引用 93 次
- When Benchmarks are Targets: Revealing the Sensitivity of Large Language Model LeaderboardsNorah A. Alzahrani, Hisham Abdullah Alyahya, Yazeed Alnumay, Sultan Alrashed 等ACL 2024 · 被引用 13 次
- Persona-Augmented Benchmarking: Evaluating LLMs Across Diverse Writing StylesKimberly Le Truong, Riccardo Fogliato, Hoda Heidari, Steven WuEMNLP 2025 · 被引用 2 次
- Active Evaluation Acquisition for Efficient LLM BenchmarkingYang Li, Jie Ma, Miguel Ballesteros, Yassine Benajiba 等ICML 2025
- Prompting is not a substitute for probability measurements in large language modelsJennifer Hu, Roger LevyEMNLP 2023 · 被引用 31 次
