Parallel Sampling via Autospeculation
Nima Anari, Carlo Baronio, CJ Chen, Alireza Haqi, Frederic Koehler, Anqi Li, Thuy-Duong Vuong
摘要
We present parallel algorithms to accelerate sampling via counting in two settings: any-order autoregressive models and denoising diffusion models. An any-order autoregressive model accesses a target distribution 𝜇 on [𝑞] 𝑛 through an oracle that provides conditional marginals, while a denoising diffusion model accesses a target distribution 𝜇 on ℝ 𝑛 through an oracle that provides conditional means under Gaussian noise. Standard sequential sampling algorithms require 𝑂(𝑛) time to produce a sample from 𝜇 in either setting. We show that, by issuing oracle calls in parallel, the expected sampling time can be reduced to 𝑂(𝑛 1/2 ). This improves the previous 𝑂(𝑛 2/3 ) bound for any-order autoregressive models and yields the first parallel speedup for diffusion models in the high-accuracy regime, under the relatively mild assumption that the support of 𝜇 is bounded.
We introduce a novel technique to obtain our results: speculative rejection sampling. This technique leverages an auxiliary "speculative" distribution 𝜈 that approximates 𝜇 to accelerate sampling. Our technique is inspired by the well-studied "speculative decoding" techniques popular in large language models, but differs in key ways. Firstly, we use "autospeculation," namely we build the speculation 𝜈 out of the same oracle that defines 𝜇. In contrast, speculative decoding typically requires a separate, faster, but potentially less accurate "draft" model 𝜈. Secondly, the key differentiating factor in our technique is that we make and accept speculations at a "sequence" level rather than at the level of single (or a few) steps. This last fact is key to unlocking our parallel runtime of 𝑂(𝑛 1/2 ).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Nearly d-Linear Convergence Bounds for Diffusion Models via Stochastic LocalizationJoe Benton, Valentin De Bortoli, Arnaud Doucet, George DeligiannidisICLR 2024 · 被引用 203 次
- The probability flow ODE is provably fastSitan Chen, Sinho Chewi, Holden Lee, Yuanzhi Li 等NeurIPS 2023 · 被引用 179 次
相关 Paper
- Diffusion Models are Secretly Exchangeable: Parallelizing DDPMs via Auto SpeculationHengyuan Hu, Aniket Das, Dorsa Sadigh, Nima AnariICML 2025
- Accelerated Diffusion Models via Speculative SamplingValentin De Bortoli, Alexandre Galashov, Arthur Gretton, Arnaud DoucetICML 2025
- Parallel Sampling via CountingNima Anari, Ruiquan Gao, Aviad RubinsteinSTOC 2024 · 被引用 2 次
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 被引用 114 次
- DFlash: Block Diffusion for Flash Speculative DecodingJian Chen, Yesheng Liang, Zhijian LiuICML 2026
