LiveBench: A Challenging, Contamination-Limited LLM Benchmark
Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Benjamin Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, Shubh-Agrawal, Sandeep Singh Sandha
摘要
Test set contamination, wherein test data from a benchmark ends up in a newer model's training set, is a well-documented obstacle for fair LLM evaluation and can quickly render benchmarks obsolete. To mitigate this, many recent benchmarks crowdsource new prompts and evaluations from human or LLM judges; however, these can introduce significant biases, and break down when scoring hard questions. In this work, we introduce a new benchmark for LLMs designed to be resistant to both test set contamination and the pitfalls of LLM judging and human crowdsourcing. We release LiveBench, the first benchmark that (1) contains frequently-updated questions from recent information sources, (2) scores answers automatically according to objective ground-truth values, and (3) contains a wide variety of challenging tasks, spanning math, coding, reasoning, language, instruction following, and data analysis. To achieve this, LiveBench contains questions that are based on recently-released math competitions, arXiv papers, news articles, and datasets, and it contains harder, contamination-limited versions of tasks from previous benchmarks such as Big-Bench Hard, AMPS, and IFEval. We evaluate many prominent closed-source models, as well as dozens of open-source models ranging from 0.5B to 405B in size. LiveBench is difficult, with top models achieving below 70% accuracy. We release all questions, code, and model answers. Questions are added and updated on a monthly basis, and we release new tasks and harder versions of tasks over time so that LiveBench can distinguish between the capabilities of LLMs as they improve in the future. We welcome community engagement and collaboration for expanding the benchmark tasks and models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper42
- GEPA: Reflective Prompt Evolution Can Outperform Reinforcement LearningLakshya A. Agrawal, Shangyin Tan, Dilara Soylu, Noah Ziems 等ICLR 2026 · 被引用 466 次
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye 等ICLR 2026 · 被引用 279 次
- FutureX: An Advanced Live Benchmark for LLM Agents in Future PredictionZhiyuan Zeng, Jiashuo Liu, Siyuan Chen, Tianci He 等ICLR 2026 · 被引用 51 次
- Model Merging in Pre-training of Large Language ModelsYunshui Li, Yiyuan Ma, Shen Yan, Chaoyi Zhang 等NeurIPS 2025 · 被引用 40 次
- Learning to Orchestrate Agents in Natural Language with the ConductorStefan Nielsen, Edoardo Cetin, Peter Schwendeman, Qi Sun 等ICLR 2026 · 被引用 22 次
它引用的顶会 Paper12
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo 等NeurIPS 2022 · 被引用 8,168 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos 等ICML 2024 · 被引用 1,212 次
相关 Paper
- LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for CodeNaman Jain, King Han, Alex Gu, Wen-Ding Li 等ICLR 2025
- LiveOIBench: Can Large Language Models Outperform Human Contestants in Informatics Olympiads?Kaijian Zou, Feiyang Xiong, Yunxiang Zhang, Xinliang Frederick Zhang 等ICML 2026 · 被引用 3 次
- AntiLeakBench: Preventing Data Contamination by Automatically Constructing Benchmarks with Updated Real-World KnowledgeXiaobao Wu, Liangming Pan, Yuxi Xie, Ruiwen Zhou 等ACL 2025 · 被引用 35 次
- Test of Time: Rethinking Temporal Signal of Benchmark ContaminationTerry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter 等ACL 2026 · 被引用 3 次
- MMLU-CF: A Contamination-free Multi-task Language Understanding BenchmarkQihao Zhao, Yangyu Huang, Tengchao Lv, Lei Cui 等ACL 2025 · 被引用 33 次
