Lune

ICLR2026顶会

Guided Speculative Inference for Efficient Test-Time Alignment of LLMs

Jonathan Geuter, Youssef Mroueh, David Alvarez-Melis

2026年份
12被引次数
6顶会引用

摘要

We propose Guided Speculative Inference (GSI), a novel algorithm for efficient reward-guided decoding in large language models. GSI combines soft best-of-nn test-time scaling with a reward model r(x,y)r(x,y) and speculative samples from a small auxiliary model πS(y∣x)\pi_S(y\mid x). We provably approximate both the optimal tilted policy πβ,B(y∣x)∝πB(y∣x)exp⁡(β r(x,y))\pi_{\beta,B}(y\mid x) \propto \pi_B(y\mid x)\exp(\beta\,r(x,y)) of soft best-of-nn under the base model πB\pi_B, as well as the expected reward under the optimal policy. In experiments on reasoning benchmarks (MATH500, OlympiadBench, Minerva Math, MMLU-STEM, GSM8K) and across different model families, our method achieves higher accuracy than standard soft best-of-nn with πS\pi_S and reward-guided speculative decoding (Liao et al., 2025), and in certain settings even outperforms soft best-of-nn with πB\pi_B, while reducing end-to-end latency by up to 28%28\%. The code is available at https://github.com/j-geuter/GSI .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper6

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖