Encoder-Decoder Diffusion Language Models for Efficient Training and Inference
Marianne Arriola, Yair Schiff, Hao Phung, Aaron Gokaslan, Volodymyr Kuleshov
摘要
Discrete diffusion models enable parallel token sampling for faster inference than autoregressive approaches. However, prior diffusion models use a decoder-only architecture, which requires sampling algorithms that invoke the full network at every denoising step and incur high computational cost. Our key insight is that discrete diffusion models perform two types of computation: 1) representing clean tokens and 2) denoising corrupted tokens, which enables us to use separate modules for each task. We propose an encoder-decoder architecture to accelerate discrete diffusion inference, which relies on an encoder to represent clean tokens and a lightweight decoder to iteratively refine a noised sequence. We also show that this architecture enables faster training of block diffusion models, which partition sequences into blocks for better quality and are commonly used in diffusion language model inference. We introduce a framework for Efficient Encoder-Decoder Diffusion (E2D2), consisting of an architecture with specialized training and sampling algorithms, and we show that E2D2 achieves superior trade-offs between generation quality and inference throughput on summarization, translation, and mathematical reasoning tasks. We provide the code 1 , model weights, and blog post on the project page: https://m-arriola.com/e2d2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Learning Unmasking Policies for Diffusion Language ModelsMetod Jazbec, Theo X. Olausson, Louis Béthune, Pierre Ablin 等ICML 2026 · 被引用 24 次
- Stop Training for the Worst: Progressive Unmasking Accelerates Masked Diffusion TrainingJaeyeon Kim, Jonathan Geuter, David Alvarez-Melis, Sham Kakade 等ICML 2026 · 被引用 8 次
- Causality in Video Diffusers is Separable from DenoisingXingjian Bai, Guande He, Zhengqi Li, Eli Shechtman 等CVPR 2026 · 被引用 3 次
- FlashBlock: Attention Caching for Efficient Long-Context Block DiffusionZhuokun Chen, Jianfei Cai, Bohan ZhuangICML 2026 · 被引用 2 次
- Locally Coherent Parallel Decoding in Diffusion Language ModelsMichael Hersche, Nicolas Menet, Ronan Tanios, Abbas RahimiICML 2026 · 被引用 1 次
它引用的顶会 Paper31
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
相关 Paper
- Set Diffusion: Interpolating Token Orderings between Autoregression and Diffusion for Fast and Flexible DecodingMarianne Arriola, Volodymyr KuleshovICML 2026 · 被引用 2 次
- Block Diffusion: Interpolating Between Autoregressive and Diffusion Language ModelsMarianne Arriola, Aaron Gokaslan, Justin T. Chiu, Zhihan Yang 等ICLR 2025
- AR-Diffusion: Auto-Regressive Diffusion Model for Text GenerationTong Wu, Zhihao Fan, Xiao Liu, Hai-Tao Zheng 等NeurIPS 2023 · 被引用 170 次
- Fast-dLLM v2: Efficient Block-Diffusion LLMChengyue Wu, Hao Zhang, Shuchen Xue, Shizhe Diao 等ICLR 2026 · 被引用 132 次
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided DiffusionZhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah 等ICLR 2026 · 被引用 56 次
