Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
Marianne Arriola, Yair Schiff, Hao Phung, Aaron Gokaslan, Volodymyr Kuleshov
Abstract
Discrete diffusion models enable parallel token sampling for faster inference than autoregressive approaches. However, prior diffusion models use a decoder-only architecture, which requires sampling algorithms that invoke the full network at every denoising step and incur high computational cost. Our key insight is that discrete diffusion models perform two types of computation: 1) representing clean tokens and 2) denoising corrupted tokens, which enables us to use separate modules for each task. We propose an encoder-decoder architecture to accelerate discrete diffusion inference, which relies on an encoder to represent clean tokens and a lightweight decoder to iteratively refine a noised sequence. We also show that this architecture enables faster training of block diffusion models, which partition sequences into blocks for better quality and are commonly used in diffusion language model inference. We introduce a framework for Efficient Encoder-Decoder Diffusion (E2D2), consisting of an architecture with specialized training and sampling algorithms, and we show that E2D2 achieves superior trade-offs between generation quality and inference throughput on summarization, translation, and mathematical reasoning tasks. We provide the code 1 , model weights, and blog post on the project page: https://m-arriola.com/e2d2.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30c10010-1dfd-4653-89dd-35462986c08eCited by top-tier papers5
- Learning Unmasking Policies for Diffusion Language ModelsMetod Jazbec, Theo X. Olausson, Louis Béthune, Pierre Ablin et al.ICML 2026 · 24 citations
- Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion TrainingJaeyeon Kim, Jonathan Geuter, David Alvarez-Melis, Sham Kakade et al.ICML 2026 · 8 citations
- Causality in Video Diffusers is Separable from DenoisingXingjian Bai, Guande He, Zhengqi Li, Eli Shechtman et al.CVPR 2026 · 3 citations
- FlashBlock: Attention Caching for Efficient Long-Context Block DiffusionZhuokun Chen, Jianfei Cai, Bohan ZhuangICML 2026 · 2 citations
- Locally Coherent Parallel Decoding in Diffusion Language ModelsMichael Hersche, Nicolas Menet, Ronan Tanios, Abbas RahimiICML 2026 · 1 citation
Builds on31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
Related papers
- Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible DecodingMarianne Arriola, Volodymyr KuleshovICML 2026 · 2 citations
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language ModelsMarianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang et al.ICLR 2025
- AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationTong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng et al.NeurIPS 2023 · 170 citations
- Fast-dLLM v2: Efficient Block-Diffusion LLMChengyue Wu, Hao Zhang, Shuchen Xue, Shizhe Diao et al.ICLR 2026 · 132 citations
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided DiffusionZhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah et al.ICLR 2026 · 56 citations
