Characterizing, Evaluating, and Optimizing Complex Reasoning
Haoran Zhang, Yafu Li, Zhi Wang, Zhilin Wang, Shunkai Zhang, Xiaoye Qu, Yu Cheng
摘要
Large Reasoning Models (LRMs) increasingly rely on reasoning traces with complex internal structures. However, existing work lacks a unified answer to three fundamental questions: (1) what defines high-quality reasoning, (2) how to reliably evaluate long, implicitly structured reasoning traces, and (3) how to use such evaluation signals for reasoning optimization. To address these challenges, we provide a unified perspective. (1) We introduce the ME principle to characterize reasoning quality along macro- and micro-level concerning efficiency and effectiveness. (2) Built on this principle, we model reasoning traces as directed acyclic graphs (DAGs) and develop a DAG-based pairwise evaluation method, capturing complex reasoning structures. (3) Based on this method, we construct the TRM-Preference dataset and train a Thinking Reward Model (TRM) to evaluate reasoning quality at scale. Experiments show that thinking rewards serve as an effective optimization signal. At test time, selecting better reasoning leads to better outcomes (up to 19.3% gain), and during RL training, thinking rewards enhance reasoning and performance (up to 3.9% gain) across diverse tasks. Code and data are available at https://github.com/Simplified-Reasoning/TRM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper17
- Self-Refine: Iterative Refinement with Self-FeedbackAman Madaan, Niket Tandon, Prakhar Gupta, Skyler Hallinan 等NeurIPS 2023 · 被引用 4,972 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Understanding the Effects of RLHF on LLM Generalisation and DiversityRobert Kirk, Ishita Mediratta, Christoforos Nalmpantis, Jelena Luketina 等ICLR 2024 · 被引用 332 次
- Learning to Reason under Off-Policy GuidanceJianhao Yan, Yafu Li, Zican Hu, Zhi Wang 等NeurIPS 2025 · 被引用 310 次
- Reinforcement Learning with Verifiable Rewards Implicitly Incentivizes Correct Reasoning in Base LLMsXumeng Wen, Zihan Liu, Shun Zheng, Shengyu Ye 等ICLR 2026 · 被引用 279 次
相关 Paper
- Reasoning Structure of Large Language ModelsFrédéric Berdoz, Luca Lanzendörfer, Fabian Farestam, Roger WattenhoferICML 2026
- Think-RM: Enabling Long-Horizon Reasoning in Generative Reward ModelsIlgee Hong, Changlong Yu, Liang Qiu, Weixiang Yan 等NeurIPS 2025 · 被引用 15 次
- Reward Reasoning ModelsJiaxin Guo, Zewen Chi, Li Dong, Qingxiu Dong 等NeurIPS 2025 · 被引用 14 次
- Smarter Not Harder: Generative Process Evaluation with Intrinsic-Signal Driving and Ability‑Adaptive Reward ShapingTao He, Rongchuan Mu, Lizi Liao, Yixin Cao 等ICLR 2026
- Unveiling Chain of Step Reasoning for Vision-Language Models with Fine-grained RewardsHonghao Chen, Xingzhou Lou, Xiaokun Feng, Kaiqi Huang 等NeurIPS 2025 · 被引用 7 次
