Fast-dLLM v2: Efficient Block-Diffusion LLM
Chengyue Wu, Hao Zhang, Shuchen Xue, Shizhe Diao, Yonggan Fu, Zhijian Liu, Pavlo Molchanov, Ping Luo, Song Han, Enze Xie
Abstract
Autoregressive (AR) large language models (LLMs) have achieved remarkable performance across a wide range of natural language tasks, yet their inherent sequential decoding limits inference efficiency. In this work, we propose Fast-dLLM v2, a carefully designed block diffusion language model (dLLM) that efficiently adapts pretrained AR models into dLLMs for parallel text generation-requiring only ∼1B tokens of fine-tuning. This represents a 500× reduction in training data compared to full-attention diffusion LLMs such as Dream (580B tokens), while preserving the original model's performance. Our approach introduces a novel training recipe that combines a block diffusion mechanism with a complementary attention mask, enabling blockwise bidirectional context modeling without sacrificing AR training objectives. To further accelerate decoding, we design a hierarchical caching mechanism: a block-level cache that stores historical context representations across blocks, and a sub-block cache that enables efficient parallel generation within partially decoded blocks. Coupled with our parallel decoding pipeline, Fast-dLLM v2 achieves up to 2.5× speedup over standard AR decoding without compromising generation quality. Extensive experiments across diverse benchmarks demonstrate that Fast-dLLM v2 matches or surpasses AR baselines in accuracy, while delivering state-of-the-art efficiency among dLLMs-marking a significant step toward the practical deployment of fast and accurate LLMs. Code and model will be publicly released. Links: Github Code | Project Page On the other hand, diffusion-based language models (dLLMs) (Google DeepMind, 2025; Inception Labs, 2025; Zhu
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2558bc19-26cd-412b-bbc4-027c046e68c5Cited by top-tier papers22
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory DistillationYu-Yang Qian, Junda Su, Lanxiang Hu, Peiyuan Zhang et al.ICML 2026 · 33 citations
- Encoder-Decoder Diffusion Language Models for Efficient Training and InferenceMarianne Arriola, Yair Schiff, Hao Phung, Aaron Gokaslan et al.NeurIPS 2025 · 21 citations
- TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model AccelerationLINYE WEI, Zixiang Luo, Pingzhi Tang, Meng LiICML 2026 · 7 citations
- Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete DiffusionLijiang Li, zuwei long, Yunhang Shen, Heting Gao et al.ICML 2026 · 7 citations
- Improving Sampling for Masked Diffusion Models via Information GainKaisen Yang, Jayden Teoh, Kaicheng Yang, Yitong Zhang et al.ICML 2026 · 6 citations
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth et al.NeurIPS 2022 · 496 citations
Related papers
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li et al.ICLR 2026 · 6 citations
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion ForcingXu Wang, Chenkai Xu, Yijie Jin, Jiachun Jin et al.ICLR 2026 · 116 citations
- Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in SpeedYonggan Fu, Lexington Whalen, Zhifan Ye, Xin Dong et al.ICML 2026 · 22 citations
- dCache: Accelerating Diffusion-Based LLMs via Dual Adaptive CachingYuchu Jiang, Yue Cai, Xiangzhong Luo, Jiale Fu et al.ICLR 2026 · 16 citations
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel DecodingChengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu et al.ICLR 2026 · 428 citations
