Parallel Sampling via Autospeculation
Nima Anari, Carlo Baronio, CJ Chen, Alireza Haqi, Frederic Koehler, Anqi Li, Thuy-Duong Vuong
Abstract
We present parallel algorithms to accelerate sampling via counting in two settings: any-order autoregressive models and denoising diffusion models. An any-order autoregressive model accesses a target distribution ๐ on [๐] ๐ through an oracle that provides conditional marginals, while a denoising diffusion model accesses a target distribution ๐ on โ ๐ through an oracle that provides conditional means under Gaussian noise. Standard sequential sampling algorithms require ๐(๐) time to produce a sample from ๐ in either setting. We show that, by issuing oracle calls in parallel, the expected sampling time can be reduced to ๐(๐ 1/2 ). This improves the previous ๐(๐ 2/3 ) bound for any-order autoregressive models and yields the first parallel speedup for diffusion models in the high-accuracy regime, under the relatively mild assumption that the support of ๐ is bounded.
We introduce a novel technique to obtain our results: speculative rejection sampling. This technique leverages an auxiliary "speculative" distribution ๐ that approximates ๐ to accelerate sampling. Our technique is inspired by the well-studied "speculative decoding" techniques popular in large language models, but differs in key ways. Firstly, we use "autospeculation," namely we build the speculation ๐ out of the same oracle that defines ๐. In contrast, speculative decoding typically requires a separate, faster, but potentially less accurate "draft" model ๐. Secondly, the key differentiating factor in our technique is that we make and accept speculations at a "sequence" level rather than at the level of single (or a few) steps. This last fact is key to unlocking our parallel runtime of ๐(๐ 1/2 ).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 3005d551-d38c-4363-aad5-804adfa15e21Cited by top-tier papers1
Ask how each one uses itBuilds on15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 ยท 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 ยท 35,902 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 ยท 1,270 citations
- Nearly d-Linear Convergence Bounds for Diffusion Models via Stochastic LocalizationJoe Benton, Valentin De Bortoli, Arnaud Doucet, George DeligiannidisICLR 2024 ยท 203 citations
- The probability flow ODE is provably fastSitan Chen, Sinho Chewi, Holden Lee, Yuanzhi Li et al.NeurIPS 2023 ยท 179 citations
Related papers
- Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto SpeculationHengyuan Hu, Aniket Das, Dorsa Sadigh, Nima AnariICML 2025
- Accelerated Diffusion Models via Speculative SamplingValentin De Bortoli, Alexandre Galashov, Arthur Gretton, Arnaud DoucetICML 2025
- Parallel Sampling via CountingNima Anari, Ruiquan Gao, Aviad RubinsteinSTOC 2024 ยท 2 citations
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 ยท 114 citations
- DFlash: Block Diffusion for Flash Speculative DecodingJian Chen, Yesheng Liang, Zhijian LiuICML 2026
