Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution
Aaron Lou, Chenlin Meng, Stefano Ermon
Abstract
Despite their groundbreaking performance for many generative modeling tasks, diffusion models have fallen short on discrete data domains such as natural language. Crucially, standard diffusion models rely on the well-established theory of score matching, but efforts to generalize this to discrete structures have not yielded the same empirical gains. In this work, we bridge this gap by proposing score entropy, a novel loss that naturally extends score matching to discrete spaces, integrates seamlessly to build discrete diffusion models, and significantly boosts performance. Experimentally, we test our Score Entropy Discrete Diffusion models (SEDD) on standard language modeling tasks. For comparable model sizes, SEDD beats existing language diffusion paradigms (reducing perplexity by -%) and is competitive with autoregressive models, in particular outperforming GPT-2. Furthermore, compared to autoregressive mdoels, SEDD generates faithful text without requiring distribution annealing techniques like temperature scaling (around - better generative perplexity than un-annealed GPT-2), can trade compute and quality (similar quality with fewer network evaluations), and enables controllable infilling (matching nucleus sampling quality while enabling other strategies besides left to right prompting).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext aff29af7-0ad2-45aa-a033-487f41eb9b1dCited by top-tier papers278
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Remasking Discrete Diffusion Models with Inference-Time ScalingGuanghan Wang, Yair Schiff, Subham S. Sahoo, Volodymyr KuleshovNeurIPS 2025 · 199 citations
- d1: Scaling Reasoning in Diffusion Large Language Models via Reinforcement LearningSiyan Zhao, Devaansh Gupta, Qinqing Zheng, Aditya GroverNeurIPS 2025 · 191 citations
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingZhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen et al.ICML 2026 · 156 citations
- LLaDA-V: Large Language Diffusion Models with Visual Instruction TuningZebin You, Shen Nie, Xiaolu Zhang, JUN ZHOU et al.CVPR 2026 · 154 citations
Builds on22
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
Related papers
- Efficient Perplexity Bound and Ratio Matching in Discrete Diffusion Language ModelsEtrit Haxholli, Yeti Ziya Gurbuz, Ogul Can, Eli WaxmanICLR 2025
- Fine-Tuning Discrete Diffusion Models with Policy Gradient MethodsOussama Zekri, Nicolas BoulléNeurIPS 2025 · 42 citations
- SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular ControlXiaochuang Han, Sachin Kumar, Yulia TsvetkovACL 2023 · 23 citations
- Target Concrete Score Matching: A Holistic Framework for Discrete DiffusionRuixiang Zhang, Shuangfei Zhai, Yizhe Zhang, James Thornton et al.ICML 2025
- A Cheaper and Better Diffusion Language Model with Soft-Masked NoiseJiaao Chen, Aston Zhang, Mu Li, Alex Smola et al.EMNLP 2023 · 16 citations
