Guided Speculative Inference for Efficient Test-Time Alignment of LLMs
Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis
Abstract
We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of- test-time scaling with a reward model and speculative samples from a small auxiliary model . We provably approximate both the optimal tilted policy of soft best-of- under the base model , as well as the expected reward under the optimal policy. In experiments on reasoning benchmarks (MATH500, OlympiadBench, Minerva Math, MMLU-STEM, GSM8K) and across different model families, our method achieves higher accuracy than standard soft best-of- with and reward-guided speculative decoding (Liao et al., 2025), and in certain settings even outperforms soft best-of- with , while reducing end-to-end latency by up to . The code is available at https://github.com/j-geuter/GSI .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bb7dc5b6-da10-4d84-a7b0-803bb505354fCited by top-tier papers6
- Best-of-N through the Smoothing Lens: KL Divergence and Regret AnalysisGholamali Aminian, Idan Shenfeld, Amir R. Asadi, Ahmad Beirami et al.ICLR 2026 · 16 citations
- Taming Imperfect Process Verifiers: A Sampling Perspective on BacktrackingDhruv Rohatgi, Abhishek Shetty, Donya Saless, Yuchen Li et al.ICLR 2026 · 15 citations
- Safety Alignment of Large Language Models via Contrasting Safe and Harmful DistributionsXiaoyun Zhang, Zhengyue Zhao, Wenxuan Shi, Kaidi Xu et al.AAAI 2026 · 4 citations
- GIFT: Bootstrapping Image-to-CAD Program Synthesis via Geometric FeedbackGiorgio Giannone, Anna Doris, Amin Nobari, Kai Xu et al.ICML 2026 · 3 citations
- Mitigating Premature Exploitation in Particle-based Monte Carlo for Inference-Time ScalingGiorgio Giannone, Guangxuan Xu, Nikhil Nayak, Rohan Awhad et al.ICML 2026 · 2 citations
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- Let's Verify Step by StepHunter Lightman, Vineet Kosaraju, Yuri Burda, Harrison Edwards et al.ICLR 2024 · 3,045 citations
- Solving Quantitative Reasoning Problems with Language ModelsAitor Lewkowycz, Anders Andreassen, David Dohan, Ethan Dyer et al.NeurIPS 2022 · 2,039 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
Related papers
- Reward-Guided Speculative Decoding for Efficient LLM ReasoningBaohao Liao, Yuhui Xu, Hanze Dong, Junnan Li et al.ICML 2025
- Accelerated Test-Time Scaling with Model-Free Speculative SamplingWoomin Song, Saket Dingliwal, Sai Muralidhar Jayanthi, Bhavana Ganesh et al.EMNLP 2025
- ATTS: Asynchronous Test-Time Scaling via Conformal PredictionJing Xiong, Qiujiang Chen, Fanghua Ye, Zhongwei Wan et al.ICLR 2026 · 8 citations
- Scaling Up, Speeding Up: A Benchmark of Speculative Decoding for Efficient LLM Test-Time ScalingShengyin Sun, Yiming Li, Xing Li, Yingzhao Lian et al.ICLR 2026 · 6 citations
- A Drop-In Solution for On-the-Fly Adaptation of Speculative Decoding in Large Language ModelsJiesong Liu, Brian Park, Xipeng ShenACL 2025 · 2 citations
