Accelerated Diffusion Models via Speculative Sampling
Valentin De Bortoli, Alexandre Galashov, Arthur Gretton, Arnaud Doucet
摘要
Speculative sampling is a popular technique for accelerating inference in Large Language Models by generating candidate tokens using a fast draft model and then accepting or rejecting them based on the target model's distribution. While speculative sampling was previously limited to discrete sequences, we extend it to diffusion models, which generate samples via continuous, vectorvalued Markov chains. In this context, the target model is a high-quality but computationally expensive diffusion model. We propose various drafting strategies, including a simple and effective approach that does not require training a draft model and is applicable out-of-the-box to any diffusion model. We demonstrate significant generation speedup on various diffusion models, halving the number of function evaluations while generating exact samples from the target model. Finally, we also show how this procedure can be used to accelerate Langevin diffusions to sample unnormalized distributions. Motivation Denoising diffusion models (DDMs), introduced by Sohl-Dickstein et al. ( 2015 ) and further developed by Ho et al. (2020) and Song et al. (2021), are generative models exhibiting state-of-the-art performance in a wide variety of domains. The core concept behind DDMs is the progressive transformation of a data distribution into a Gaussian distribution through the addition of noise. Sample generation is achieved by simulating an approximation of the time-reversal of this noising process. This requires multiple evaluations of a neural network that approximates the scores of the noising process, and typically involves simulating a Markov chain over hundreds of steps. Since sample generation is computationally expensive, sev-* Equal contribution 1 Google DeepMind.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Accelerated Sampling from Masked Diffusion Models via Entropy Bounded UnmaskingHeli Ben-Hamu, Itai Gat, Daniel Severo, Niklas Nolte 等NeurIPS 2025 · 被引用 131 次
- How to build a consistency model: Learning flow maps via self-distillationNicholas M. Boffi, Michael S. Albergo, Eric Vanden-EijndenNeurIPS 2025 · 被引用 111 次
- Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsYinuo Ren, Haoxuan Chen, Yuchen Zhu, Wei Guo 等NeurIPS 2025 · 被引用 51 次
- Inductive Generative Recommendation via Retrieval-based SpeculationYijie Ding, Jiacheng Li, Julian J. McAuley, Yupeng HouAAAI 2026 · 被引用 19 次
- Generalised Flow Maps for Few-Step Generative Modelling on Riemannian ManifoldsOscar Davis, Michael S. Albergo, Nicholas M. Boffi, Michael M. Bronstein 等ICLR 2026 · 被引用 12 次
它引用的顶会 Paper21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
相关 Paper
- DFlash: Block Diffusion for Flash Speculative DecodingJian Chen, Yesheng Liang, Zhijian LiuICML 2026
- Parallel Sampling via AutospeculationNima Anari, Carlo Baronio, CJ Chen, Alireza Haqi 等STOC 2026 · 被引用 5 次
- Speculative Decoding with CTC-based Draft Model for LLM Inference AccelerationZhuofan Wen, Shangtong Gui, Yang FengNeurIPS 2024 · 被引用 19 次
- Speculative Sampling For Faster Molecular DynamicsArthur Kosmala, Stephan Günnemann, Meng Gao, Brandon WoodICML 2026
- A Theoretical Perspective for Speculative Decoding AlgorithmMing Yin, Minshuo Chen, Kaixuan Huang, Mengdi WangNeurIPS 2024 · 被引用 36 次
