ReEfBench: Quantifying the Reasoning Efficiency of LLMs
Zhizhang Fu, Yuancheng Gu, Chenkai Hu, Hanmeng Liu, Yue Zhang
Abstract
Test-time scaling has enabled Large Language Models (LLMs) to tackle complex reasoning, yet the limitations of current Chain-of-Thought (CoT) evaluation obscures whether performance gains stem from genuine reasoning or mere verbosity. To address this, (1) we propose a novel neuro-symbolic framework for the non-intrusive, comprehensive process-centric evaluation of reasoning. (2) Through this lens, we identify four distinct behavioral prototypes and diagnose the failure modes. (3) We examine the impact of inference mode, training strategy, and model scale. Our analysis reveals that extended token generation is not a prerequisite for deep reasoning. Furthermore, we reveal critical constraints: mixing long and short CoT data in training risks in premature saturation and collapse, while distillation into smaller models captures behavioral length but fails to replicate logical efficacy due to intrinsic capacity limits.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 623f608a-60ee-4e75-9a4f-3d9e3250aff5Builds on10
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Testing the General Deductive Reasoning Capacity of Large Language Models Using OOD ExamplesAbulhair Saparov, Richard Yuanzhe Pang, Vishakh Padmakumar, Nitish Joshi et al.NeurIPS 2023 · 145 citations
- Language Models Are Greedy Reasoners: A Systematic Formal Analysis of Chain-of-ThoughtAbulhair Saparov, He HeICLR 2023 · 38 citations
- LINC: A Neurosymbolic Approach for Logical Reasoning by Combining Language Models with First-Order Logic ProversTheo Olausson, Alex Gu, Benjamin Lipkin, Cedegao E. Zhang et al.EMNLP 2023 · 37 citations
- ROSCOE: A Suite of Metrics for Scoring Step-by-Step ReasoningOlga Golovneva, Moya Chen, Spencer Poff, Martin Corredor et al.ICLR 2023 · 28 citations
Related papers
- Mapping the Minds of LLMs: A Graph-Based Analysis of Reasoning LLMsZhen Xiong, Yujun Cai, Zhecheng Li, Yiwei WangEMNLP 2025
- What Characterizes Effective Reasoning? Revisiting Length, Review, and Structure of CoTYunzhen Feng, Julia Kempe, Cheng Zhang, Parag Jain et al.ICML 2026
- When More is Less: Understanding Chain-of-Thought Length in LLMsYuyang Wu, Yifei Wang, Ziyu Ye, Tianqi Du et al.ICLR 2026 · 225 citations
- Too Long, Do Re-weighting for Efficient LLM Reasoning CompressionZhong-Zhi Li, Xiao Liang, Zihao Tang, Lei Ji et al.ACL 2026 · 5 citations
- Long-Context Reasoning Through Proxy-Based Chain-of-Thought TuningMiao Li, Irina Saparina, Alexander Gurung, Mirella LapataACL 2026
