Scaling Speculative Decoding with Lookahead Reasoning
Yichao Fu, Rui Ge, Zelei Shao, Zhijie Deng, Hao Zhang
摘要
Reasoning models excel by generating long chain-of-thoughts, but decoding the resulting thousands of tokens is slow. Token-level speculative decoding (SD) helps, but its benefit is capped, because the chance that an entire γ-token guess is correct falls exponentially as γ grows. This means allocating more compute for longer token drafts faces an algorithmic ceiling -making the speedup modest and hardware-agnostic. We raise this ceiling with LOOKAHEAD REASONING, which exploits a second, step-level layer of parallelism. Our key insight is that reasoning models generate step-by-step, and each step needs only to be semantically correct, not exact token matching. In LOOKAHEAD REASONING, a lightweight draft model proposes several future steps; the target model expands each proposal in one batched pass, and a verifier keeps semantically correct steps while letting the target regenerate any that fail. Token-level SD still operates within each reasoning step, so the two layers of parallelism multiply. We show LOOKAHEAD REASONING lifts the peak speedup of SD both theoretically and empirically. Across GSM8K, AIME, and other benchmarks, LOOKAHEAD REASONING improves the speedup of SD from 1.4x to 2.1x while preserving answer quality, and its speedup scales better with additional GPU throughput. Our code is available at https://github. com/hao-ai-lab/LookaheadReasoning Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Speculative Actions: A Lossless Framework for Faster AI AgentsNaimeng Ye, Arnav Ahuja, Georgios Liargkovas, Yunan Lu 等ICLR 2026 · 被引用 9 次
- Small Drafts, Big Verdict: Information-Intensive Visual Reasoning via SpeculationYuhan Liu, Lianhui Qin, Shenji WanICLR 2026 · 被引用 6 次
- ConfSpec: Efficient Step-Level Speculative Reasoning via Confidence-Gated VerificationSiran Liu, Zane Cao, Yongchao HeACL 2026 · 被引用 3 次
- Thinking with Drafts: Speculative Temporal Reasoning for Efficient Long Video UnderstandingPengfei Hu, Meng Cao, Yingyao Wang, Yi Wang 等CVPR 2026 · 被引用 3 次
- When Drafts Evolve: Speculative Decoding Meets Online LearningYu-Yang Qian, Hao-Cong Wu, Yichao Fu, Hao Zhang 等ICML 2026 · 被引用 2 次
它引用的顶会 Paper8
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma 等NeurIPS 2022 · 被引用 22,562 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- EAGLE-3: Scaling up Inference Acceleration of Large Language Models via Training-Time TestYuhui Li, Fangyun Wei, Chao Zhang, Hongyang ZhangNeurIPS 2025 · 被引用 347 次
- Break the Sequential Dependency of LLM Inference Using Lookahead DecodingYichao Fu, Peter Bailis, Ion Stoica, Hao ZhangICML 2024 · 被引用 290 次
相关 Paper
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningRui Pan, Yinwei Dai, Zhihao Zhang, Gabriele Oliaro 等NeurIPS 2025 · 被引用 68 次
- Speculative Speculative DecodingTanishq Kumar, Tri Dao, Avner MayICLR 2026 · 被引用 15 次
- Calibrated Speculative Decoding: Frequency-Guided Candidate Selection for Efficient InferenceXuwen Zhou, Fangxin Liu, Chao Wang, Xiao Zheng 等ACL 2026
- Accelerated Test-Time Scaling with Model-Free Speculative SamplingWoomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh 等EMNLP 2025
- Taming the Long-Tail: Efficient Reasoning RL Training with Adaptive DrafterQinghao Hu, Shang Yang, Junxian Guo, Xiaozhe Yao 等ASPLOS 2026 · 被引用 1 次
