Lune

ICLR2026Top-tier venue

Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis

2026Year
12Citations
6Top-tier citations

Abstract

We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of-nn test-time scaling with a reward model r(x,y)r(x,y) and speculative samples from a small auxiliary model πS(y∣x)\pi_S(y\mid x). We provably approximate both the optimal tilted policy πβ,B(y∣x)∝πB(y∣x)exp⁡(β r(x,y))\pi_{\beta,B}(y\mid x) \propto \pi_B(y\mid x)\exp(\beta\,r(x,y)) of soft best-of-nn under the base model πB\pi_B, as well as the expected reward under the optimal policy. In experiments on reasoning benchmarks (MATH500, OlympiadBench, Minerva Math, MMLU-STEM, GSM8K) and across different model families, our method achieves higher accuracy than standard soft best-of-nn with πS\pi_S and reward-guided speculative decoding (Liao et al., 2025), and in certain settings even outperforms soft best-of-nn with πB\pi_B, while reducing end-to-end latency by up to 28%28\%. The code is available at https://github.com/j-geuter/GSI .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext bb7dc5b6-da10-4d84-a7b0-803bb505354f

Cited by top-tier papers6

Ask how each one uses it

Builds on20

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines