Benchmarking at the Edge of Comprehension
Samuele Marro, Jialin Yu, Emanuele La Malfa, Oishi Deb, Jiawei Li, Yibo Yang, Ebey Abraham, Sunando Sengupta, Eric Sommerlade, Michael Wooldridge, Phil Torr
Abstract
As frontier Large Language Models (LLMs) increasingly saturate new benchmarks shortly after they are published, benchmarking itself is at a juncture: if frontier models keep improving, it will become increasingly hard for humans to generate discriminative tasks, provide accurate ground-truth answers, or evaluate complex solutions. If benchmarking becomes infeasible, our ability to measure any progress in AI is at stake. We refer to this scenario as the post-comprehension regime . In this work, we propose Critique-Resilient Benchmarking, an adversarial framework designed to compare models even when full human understanding is infeasible. Our technique relies on the notion of critique-resilient correctness : an answer is deemed correct if no adversary has convincingly proved otherwise. Unlike standard benchmarking, humans serve as bounded verifiers and focus on localized claims, which preserves evaluation integrity beyond full comprehension of the task. Using an itemized bipartite Bradley-Terry model, we jointly rank LLMs by their ability to solve challenging tasks and to generate difficult yet solvable questions. We showcase the effectiveness of our method in the mathematical domain across eight frontier LLMs, showing that the resulting scores are stable and correlate with external capability measures. Our framework reformulates benchmarking as an adversarial generation-evaluation game in which humans serve as final adjudicators.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e6ece554-1ef2-4b0e-8ceb-339c3b34e557Builds on8
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Chatbot Arena: An Open Platform for Evaluating LLMs by Human PreferenceWei-Lin Chiang, Lianmin Zheng, Ying Sheng, Anastasios Nikolas Angelopoulos et al.ICML 2024 · 1,212 citations
- Self-Instruct: Aligning Language Models with Self-Generated InstructionsYizhong Wang, Yeganeh Kordi, Swaroop Mishra, Alisa Liu et al.ACL 2023 · 540 citations
- miniF2F: a cross-system benchmark for formal Olympiad-level mathematicsKunhao Zheng, Jesse Michael Han, Stanislas PoluICLR 2022 · 342 citations
- On scalable oversight with weak LLMs judging strong LLMsZachary Kenton, Noah Y. Siegel, János Kramár, Jonah Brown-Cohen et al.NeurIPS 2024 · 116 citations
Related papers
- Expanding the AI Evaluation Toolbox with Statistical ModelsDrew Keller, Kweku Kwegyir-Aggrey, Ryan Steed, Anita K Rao et al.ICML 2026 · 4 citations
- Pressure Reveals Character: Behavioural Alignment Evaluation at DepthNora Petrova, John BurdenICML 2026
- From Winning to Understanding: A Diagnostic Long-Horizon RTS Benchmark for LLMsJiacheng Li, Jiahui Liu, Yuqing Wang, Gaochen Cui et al.ICML 2026
- EIP: Weighted Ranking of LLMs by Quantifying Question DifficultyXingjian Hu, Ziqian Zhang, Yue Huang, Kai Zhang et al.ICLR 2026
- Potemkin Understanding in Large Language ModelsMarina Mancoridis, Bec Weeks, Keyon Vafa, Sendhil MullainathanICML 2025
