Scaling Diffusion Language Models via Adaptation from Autoregressive Models
Shansan Gong, Shivam Agarwal, Yizhe Zhang, Jiacheng Ye, Lin Zheng, Mukai Li, Chenxin An, Peilin Zhao, Wei Bi, Jiawei Han, Hao Peng, Lingpeng Kong
Abstract
Diffusion Language Models (DLMs) have emerged as a promising new paradigm for text generative modeling, potentially addressing limitations of autoregressive (AR) models. However, current DLMs have been studied at a smaller scale compared to their AR counterparts and lack fair comparison on language modeling benchmarks. Additionally, training diffusion models from scratch at scale remains challenging. Given the prevalence of open-source AR language models, we propose adapting these models to build text diffusion models. We demonstrate connections between AR and diffusion modeling objectives and introduce a simple continual pre-training approach for training diffusion models. Through systematic evaluation on language modeling, reasoning, and commonsense benchmarks, we show that we can convert AR models ranging from 127M to 7B parameters (GPT2 and LLaMA) into diffusion models DiffuGPT and DiffuLLaMA, using less than 200B tokens for training. Our experimental results reveal that these models outperform earlier DLMs and are competitive with their AR counterparts. We release a suite of DLMs (127M-355M-7B) capable of generating fluent text, performing in-context learning, filling in the middle without prompt re-ordering, and following instructions. https://github.com/HKUNLP/DiffuLLaMA
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9b87b3cb-ee71-4032-9e0f-1a792daab28dCited by top-tier papers100
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel DecodingChengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu et al.ICLR 2026 · 428 citations
- LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion ModelsFengqi Zhu, Rongzhen Wang, Shen Nie, Xiaolu Zhang et al.ACL 2026 · 229 citations
- DiffuCoder: Understanding and Improving Masked Diffusion Models for Code GenerationShansan Gong, Ruixiang Zhang, Huangjie Zheng, Jiatao Gu et al.ICLR 2026 · 198 citations
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningSiyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya GroverNeurIPS 2025 · 191 citations
Builds on47
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
Related papers
- Beyond Autoregression: Fast LLMs via Self-Distillation Through TimeJustin Deschenaux, Caglar GulcehreICLR 2025
- Likelihood-Based Diffusion Language ModelsIshaan Gulrajani, Tatsunori B. HashimotoNeurIPS 2023 · 178 citations
- Text Generation with Diffusion Language Models: A Pre-training Approach with Continuous Paragraph DenoiseZhenghao Lin, Yeyun Gong, Yelong Shen, Tong Wu et al.ICML 2023 · 107 citations
- Scaling Behavior of Discrete Diffusion Language ModelsDimitri von Rütte, Janis Fluri, Omead Pooladzandi, Bernhard Schölkopf et al.ICLR 2026 · 34 citations
- Scaling up Masked Diffusion Models on TextShen Nie, Fengqi Zhu, Chao Du, Tianyu Pang et al.ICLR 2025
