Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed
Yonggan Fu, Lexington Whalen, Zhifan Ye, Xin Dong, Shizhe Diao, Jingyu Liu, CHENGYUE WU, Hao Zhang, Enze Xie, Song Han, Maksim Khadkevich, Jan Kautz
Abstract
Diffusion language models (dLMs) have emerged as a promising paradigm enabling parallel generation, but their learning efficiency lags behind that of autoregressive (AR) language models when trained from scratch. To this end, we study AR-to-dLM conversion, which transforms pretrained AR models into efficient dLMs that excel in speed while preserving AR models' task accuracy. We achieve this by identifying limitations in the attention patterns and objectives of existing AR-to-dLM methods and then proposing methodologies and actionable insights for scalable AR-to-dLM conversion. Specifically, we first systematically compare different attention patterns and find that maintaining pretrained AR weight distributions is key to effective AR-to-dLM conversion. Accordingly, we introduce a continuous pretraining scheme with a block-wise attention pattern, which remains causal across blocks with bidirectional modeling within each block. We find that, in addition to block-wise attention's known benefit of enabling KV caching, its block-wise causality better preserves pretrained AR models' weight distributions, leading to a win-win in accuracy and efficiency. Second, to mitigate the training-test gap in mask token distributions (uniform vs. highly left-to-right), we propose a position-dependent token masking strategy that assigns higher masking probabilities to later tokens during training to better mimic test-time behavior. Leveraging this framework, we conduct extensive studies of dLMs' attention patterns, training dynamics, and other design choices. These studies lead to the Efficient-DLM model family, which outperforms state-of-the-art AR models and dLMs in accuracy-throughput trade-offs; for example, our Efficient-DLM 8B achieves +5.4%/+2.7% higher accuracy with 4.5×/2.7× higher throughput compared to Dream 7B and Qwen3 4B, respectively.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1b6ca191-e857-49ee-8e08-0a78625576f8Cited by top-tier papers6
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast InferenceAiwei Liu, Minghua He, Shaoxun Zeng, Sijun Zhang et al.ICML 2026 · 37 citations
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory DistillationYu-Yang Qian, Junda Su, Lanxiang Hu, Peiyuan Zhang et al.ICML 2026 · 33 citations
- TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model AccelerationLINYE WEI, Zixiang Luo, Pingzhi Tang, Meng LiICML 2026 · 7 citations
- When Drafts Evolve: Speculative Decoding Meets Online LearningYu-Yang Qian, Hao-Cong Wu, Yichao Fu, Hao Zhang et al.ICML 2026 · 2 citations
- Locally Coherent Parallel Decoding in Diffusion Language ModelsMichael Hersche, Nicolas Menet, Ronan Tanios, Abbas RahimiICML 2026 · 1 citation
Builds on24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
Related papers
- Fast-dLLM v2: Efficient Block-Diffusion LLMChengyue Wu, Hao Zhang, Shuchen Xue, Shizhe Diao et al.ICLR 2026 · 132 citations
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion ForcingXu Wang, Chenkai Xu, Yijie Jin, Jiachun Jin et al.ICLR 2026 · 116 citations
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li et al.ICLR 2026 · 6 citations
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky et al.ICML 2026
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingZhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen et al.ICML 2026 · 156 citations
