Diffusion Forcing: Next-token Prediction Meets Full-Sequence Diffusion
Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, Vincent Sitzmann
Abstract
This paper presents Diffusion Forcing, a new training paradigm where a diffusion model is trained to denoise a set of tokens with independent per-token noise levels. We apply Diffusion Forcing to sequence generative modeling by training a causal next-token prediction model to generate one or several future tokens without fully diffusing past ones. Our approach is shown to combine the strengths of next-token prediction models, such as variable-length generation, with the strengths of full-sequence diffusion models, such as the ability to guide sampling to desirable trajectories. Our method offers a range of additional capabilities, such as (1) rolling-out sequences of continuous tokens, such as video, with lengths past the training horizon, where baselines diverge and (2) new sampling and guiding schemes that uniquely profit from Diffusion Forcing's variable-horizon and causal architecture, and which lead to marked performance gains in decision-making and planning tasks. In addition to its empirical success, our method is proven to optimize a variational lower bound on the likelihoods of all subsequences of tokens drawn from the true joint distribution. Project website: https://boyuan.space/diffusion-forcing
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 35d7daea-963d-43a6-adac-cd87d86b5eb6Cited by top-tier papers216
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao et al.ICLR 2026 · 241 citations
- HybridVLA: Collaborative Diffusion and Autoregression in a Unified Vision-Language-Action ModelJiaming Liu, Hao Chen, Zhuoyang Liu, Pengju An et al.ICLR 2026 · 216 citations
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan et al.ICLR 2026 · 215 citations
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang et al.ICLR 2026 · 181 citations
Builds on30
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative CompressionJung Yi, Wooseok Jang, Paul Cho, Jisu Nam et al.ICML 2026
- Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion ModelXinyin Ma, Julius Berner, Chao Liu, Arash Vahdat et al.ICML 2026
- History-Guided Video DiffusionKiwhan Song, Boyuan Chen, Max Simchowitz, Yilun Du et al.ICML 2025
- FloodDiffusion: Tailored Diffusion Forcing for Streaming Motion GenerationYIYI CAI, Yuhan Wu, Kunhang Li, YOU ZHOU et al.CVPR 2026 · 14 citations
- Self-Speculative Masked DiffusionsAndrew Campbell, Valentin De Bortoli, Jiaxin Shi, Arnaud DoucetICLR 2026 · 12 citations
