SpecReason: Fast and Accurate Inference-Time Compute via Speculative Reasoning
Rui Pan, Yinwei Dai, Zhihao Zhang, Gabriele Oliaro, Zhihao Jia, Ravi Netravali
摘要
Recent advances in inference-time compute have significantly improved performance on complex tasks by generating long chains of thought (CoTs) using Large Reasoning Models (LRMs). However, this improved accuracy comes at the cost of high inference latency due to the length of generated reasoning sequences and the autoregressive nature of decoding. Our key insight in tackling these overheads is that LRM inference, and the reasoning that it embeds, is highly tolerant of approximations: complex tasks are typically broken down into simpler steps, each of which brings utility based on the semantic insight it provides for downstream steps rather than the exact tokens it generates. Accordingly, we introduce SpecReason, a system that automatically accelerates LRM inference by using a lightweight model to (speculatively) carry out simpler intermediate reasoning steps and reserving the costly base model only to assess (and potentially correct) the speculated outputs. Importantly, SpecReason's focus on exploiting the semantic flexibility of thinking tokens in preserving final-answer accuracy is complementary to prior speculation techniques, most notably speculative decoding, which demands token-level equivalence at each step. Across a variety of cross-domain reasoning benchmarks, SpecReason achieves 1.4 -3.0× speedup over vanilla LRM inference while improving accuracy by 0.4 -9.0%. Compared to speculative decoding without SpecReason, their combination yields an additional 8.8 -58.0% latency reduction. We open-source SpecReason at https://github.com/ruipeterpan/specreason.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- R2R: Efficiently Navigating Divergent Reasoning Paths with Small-Large Model Token RoutingTianyu Fu, Yi Ge, Yichen You, Enshu Liu 等NeurIPS 2025 · 被引用 32 次
- VeriThinker: Learning to Verify Makes Reasoning Model EfficientZigeng Chen, Xinyin Ma, Gongfan Fang, Ruonan Yu 等NeurIPS 2025 · 被引用 30 次
- Agentic Plan Caching: Test-Time Memory for Fast and Cost-Efficient LLM AgentsQizheng Zhang, Michael Wornow, Kunle OlukotunNeurIPS 2025 · 被引用 27 次
- Scaling Speculative Decoding with Lookahead ReasoningYichao Fu, Rui Ge, Zelei Shao, Zhijie Deng 等NeurIPS 2025 · 被引用 14 次
- TrimR: Verifier-based Training-Free Thinking Trimming for Efficient Test-Time ScalingWeizhe Lin, Xing Li 023, Zhiyuan Yang, Xiaojin Fu 等ICLR 2026 · 被引用 14 次
它引用的顶会 Paper14
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran 等NeurIPS 2023 · 被引用 5,068 次
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards 等ICLR 2024 · 被引用 3,045 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
相关 Paper
- Breaking the Reward Barrier: Accelerating Tree-of-Thought Reasoning via Speculative ExplorationShuzhang Zhong, Haochen Huang, Shengxuan Qiu, Pengfei Zuo 等OSDI 2026
- SpecExit: Accelerating Large Reasoning Model via Speculative ExitRubing Yang, Huajun Bai, Song Liu, Guanghua Yu 等ICML 2026 · 被引用 7 次
- ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated VerificationSiran Liu, Zane Cao, Yongchao HeACL 2026 · 被引用 3 次
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive DrafterQinghao Hu, Shang Yang, Junxian Guo, Xiaozhe Yao 等ASPLOS 2026 · 被引用 1 次
- System-1.5 Reasoning: Traversal in Language and Latent Spaces with Dynamic ShortcutsXiaoqiang Wang, Suyuchen Wang, Yun Zhu, Bang LiuNeurIPS 2025 · 被引用 26 次
