JudgeBoard: Benchmarking and Enhancing Small Language Models for Reasoning Evaluation
Zhenyu Bi, Gaurav Srivastava, Yang Li, Swastik Roy, Meng Lu, Morteza Ziyadi, Xuan Wang
摘要
Evaluating language models fairly is becoming harder as static benchmarks risk contamination by training data, making it unclear whether models are truly reasoning or just recalling answers. We introduce BEYONDBENCH 1 , an evaluation framework that avoids this problem by using algorithmic problem generation. Unlike traditional benchmarks that risk contamination from internet-scale training data, BEYONDBENCH creates mathematically grounded problems on the fly, ensuring each test remains fresh and uncontaminated. Our framework covers 44 algorithmic tasks with a total of 117 variations, grouped into three difficulty levels: the Easy Suite (29 tasks) for basic arithmetic and statistics, the Medium Suite (5 tasks, 49 variations) for sequence patterns and reasoning, and the Hard Suite (10 tasks, 68 variations) tackling NP-complete and constraint satisfaction problems. Each task generates problems from a combinatorial space larger than 10 15 unique instances, with solutions verified deterministically by mathematical proofs. We evaluated 101 language models, including 85 open-source and 16 closed-source models, spanning sizes from 0.5B to 141B parameters and multiple quantization schemes. All evaluations use three-fold evaluation to ensure statistical robustness. Our results show consistent reasoning deficiencies across model families, with performance degrading sharply as problem complexity increases from polynomial to exponential. In our Hard Suite evaluations, models such as Gemini-2.5-pro, Llama-3.3-70B, and Qwen2.5-72B achieved average accuracies of 56.21%, 27.16%, and 33.37%, respectively. Moreover, we observe that performance drops drastically without tool usage, with GPT-5, GPT-5-mini, and GPT-5-nano showing a decline of 16.81%, 15.86%, and 43.95% in overall accuracy without tool access. The contamination resistance of BEYONDBENCH rests on three guarantees: (i) the problem space is vastly larger than any static dataset, (ii) every instance has a deterministically verifiable solution (unique or fully enumerated), and (iii) isomorphic transformations generate semantically equivalent but syntactically new problems. BEYONDBENCH redefines reasoning evaluation through genuine algorithmic problem-solving, ensuring fair and meaningful evaluation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- PAL: Program-aided Language ModelsLuyu Gao, Aman Madaan, Shuyan Zhou, Uri Alon 等ICML 2023 · 被引用 700 次
- The Illusion of Thinking: Understanding the Strengths and Limitations of Reasoning Models via the Lens of Problem ComplexityParshin Shojaee, Iman Mirzadeh, Keivan Alizadeh-Vahid, Maxwell Horton 等NeurIPS 2025 · 被引用 507 次
- FP8 Quantization: The Power of the ExponentAndrey Kuzmin, Mart van Baalen, Yuwei Ren, Markus Nagel 等NeurIPS 2022 · 被引用 154 次
- The CLRS Algorithmic Reasoning BenchmarkPetar Velickovic, Adrià Puigdomènech Badia, David Budden, Razvan Pascanu 等ICML 2022 · 被引用 118 次
相关 Paper
- BeyondBench: Contamination-Resistant Evaluation of Reasoning in Language ModelsGaurav Srivastava, Aafiya Hussain, Zhenyu Bi, Swastik Roy 等ICLR 2026 · 被引用 4 次
- UGMathBench: A Diverse and Dynamic Benchmark for Undergraduate-Level Mathematical Reasoning with Large Language ModelsXin Xu, Jiaxin Zhang, Tianhao Chen, Zitong Chao 等ICLR 2025
- RefineBench: Evaluating Refinement Capability of Language Models via ChecklistsYoung-Jun Lee, Seungone Kim, Byung-Kwan Lee, Minkyeong Moon 等ICLR 2026 · 被引用 13 次
- BizBench: A Quantitative Reasoning Benchmark for Business and FinanceMichael Krumdick, Rik Koncel-Kedziorski, Viet Dac Lai, Varshini Reddy 等ACL 2024 · 被引用 10 次
- LiveBench: A Challenging, Contamination-Limited LLM BenchmarkColin White, Samuel Dooley, Manley Roberts, Arka Pal 等ICLR 2025
