Think while You Generate: Discrete Diffusion with Planned Denoising
Sulin Liu, Juno Nam, Andrew Campbell, Hannes Stärk, Yilun Xu, Tommi S. Jaakkola, Rafael Gómez-Bombarelli
Abstract
Discrete diffusion has achieved state-of-the-art performance, outperforming or approaching autoregressive models on standard benchmarks. In this work, we introduce Discrete Diffusion with Planned Denoising (DDPD), a novel framework that separates the generation process into two models: a planner and a denoiser. At inference time, the planner selects which positions to denoise next by identifying the most corrupted positions in need of denoising, including both initially corrupted and those requiring additional refinement. This plan-and-denoise approach enables more efficient reconstruction during generation by iteratively identifying and denoising corruptions in the optimal order. DDPD outperforms traditional denoiser-only mask diffusion methods, achieving superior results on language modeling benchmarks such as text8, OpenWebText, and token-based generation on ImageNet 256 × 256. Notably, in language modeling, DDPD significantly reduces the performance gap between diffusion-based and autoregressive methods in terms of generative perplexity. Code is available at github.com/liusulin/DDPD.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ba3efd5e-998b-45d7-83c6-92f9734a49f3Cited by top-tier papers25
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Accelerated Sampling from Masked Diffusion Models via Entropy Bounded UnmaskingHeli Ben-Hamu, Itai Gat, Daniel Severo, Niklas Nolte et al.NeurIPS 2025 · 131 citations
- Fast Solvers for Discrete Diffusion Models: Theory and Applications of High-Order AlgorithmsYinuo Ren, Haoxuan Chen, Yuchen Zhu, Wei Guo et al.NeurIPS 2025 · 51 citations
- KLASS: KL-Guided Fast Inference in Masked Diffusion ModelsSeo Hyun Kim, Sunwoo Hong, Hojung Jung, Youngrok Park et al.NeurIPS 2025 · 48 citations
- Fine-Tuning Masked Diffusion for Provable Self-CorrectionJaeyeon Kim, Seunggeun Kim, Taekyun Lee, David Pan et al.ICML 2026 · 35 citations
Builds on7
- Vector-quantized Image Modeling with Improved VQGANJiahui Yu, Xin Li, Jing Yu Koh, Han Zhang et al.ICLR 2022 · 753 citations
- Simplified and Generalized Masked Diffusion for Discrete DataJiaxin Shi, Kehang Han, Zhe Wang, Arnaud Doucet et al.NeurIPS 2024 · 693 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
- An Image is Worth 32 Tokens for Reconstruction and GenerationQihang Yu, Mark Weber, Xueqing Deng, Xiaohui Shen et al.NeurIPS 2024 · 331 citations
- Step-unrolled Denoising Autoencoders for Text GenerationNikolay Savinov, Junyoung Chung, Mikolaj Binkowski, Erich Elsen et al.ICLR 2022 · 142 citations
Related papers
- Planned DiffusionDaniel Mingyi Israel, Tian Jin, Ellie Y Cheng, Guy Van den Broeck et al.ICLR 2026 · 8 citations
- Plan for Speed: Dilated Scheduling for Masked Diffusion Language ModelsOmer Luxembourg, Haim Permuter, Eliya NachmaniICML 2026 · 30 citations
- Unifying Continuous and Discrete Text Diffusion with Non-simultaneous Diffusion ProcessesBocheng Li, Zhujin Gao, Linli XuACL 2025
- Planner Aware Path Learning in Diffusion Language Models TrainingFred Zhangzhi Peng, Zachary Bezemek, Jarrid Rector-Brooks, Shuibai Zhang et al.ICLR 2026 · 14 citations
- Flexible-length Text Infilling for Discrete Diffusion ModelsAndrew Zhang, Anushka Sivakumar, Chia-Wei Tang, Chris ThomasEMNLP 2025 · 11 citations
