Accelerated Test-Time Scaling with Model-Free Speculative Sampling
Woomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh, Jinwoo Shin, Aram Galstyan, Sravan Babu Bodapati
摘要
Language models have demonstrated remarkable capabilities in reasoning tasks through test-time scaling techniques like best-of-N sampling and tree search. However, these approaches often demand substantial computational resources, creating a critical trade-off between performance and efficiency. We introduce STAND (STochastic Adaptive N-gram Drafting), a novel model-free speculative decoding approach that exploits the inherent redundancy in reasoning trajectories to achieve significant acceleration without compromising accuracy. Our analysis shows that reasoning paths frequently reuse similar reasoning patterns, enabling efficient model-free token prediction without requiring separate draft models. By introducing stochastic drafting and preserving probabilistic information through a memory-efficient logit-based N-gram module, combined with optimized Gumbel-Top-K sampling and data-driven tree construction, STAND significantly improves token acceptance rates. Extensive evaluations across multiple models and reasoning tasks (AIME-2024, GPQA-Diamond, and LiveCodeBench) demonstrate that STAND reduces inference latency by 60-65% compared to standard autoregressive decoding while maintaining accuracy. Furthermore, STAND consistently outperforms stateof-the-art speculative decoding methods across diverse inference patterns, including singletrajectory decoding, batch decoding, and testtime tree search. As a model-free approach, STAND can be applied to any existing language model without additional training, making it a powerful plug-and-play solution for accelerating language model reasoning.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- Efficient Training-Free Multi-Token Prediction via Embedding-Space ProbingRaghavv Goel, Mukul Gagrani, Mingu Lee, Christopher LottICML 2026
它引用的顶会 Paper8
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- Medusa: Simple LLM Inference Acceleration Framework with Multiple Decoding HeadsTianle Cai, Yuhong Li, Zhengyang Geng, Hongwu Peng 等ICML 2024 · 被引用 669 次
- Escape Sky-high Cost: Early-stopping Self-Consistency for Multi-step ReasoningYiwei Li, Peiwen Yuan, Shaoxiong Feng, Boyuan Pan 等ICLR 2024 · 被引用 101 次
- SpecReason: Fast and Accurate Inference-Time Compute via Speculative ReasoningRui Pan, Yinwei Dai, Zhihao Zhang, Gabriele Oliaro 等NeurIPS 2025 · 被引用 68 次
相关 Paper
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time ScalingShengyin Sun, Yiming Li, Xing Li, Yingzhao Lian 等ICLR 2026 · 被引用 6 次
- Not-a-Bandit: Provably No-Regret Drafter Selection in Speculative Decoding for LLMsHongyi Liu, Jiaji Huang, Zhen Jia, Youngsuk Park 等ICLR 2026 · 被引用 5 次
- ATTS: Asynchronous Test-Time Scaling via Conformal PredictionJing Xiong, Qiujiang Chen, Fanghua Ye, Zhongwei Wan 等ICLR 2026 · 被引用 8 次
- Scaling Speculative Decoding with Lookahead ReasoningYichao Fu, Rui Ge, Zelei Shao, Zhijie Deng 等NeurIPS 2025 · 被引用 14 次
- CAS-Spec: Cascade Adaptive Self-Speculative Decoding for On-the-Fly Lossless Inference Acceleration of LLMsZhiyuan Ning, Jiawei Shao, Ruge Xu, Xinfei Guo 等NeurIPS 2025 · 被引用 5 次
