WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast Inference
Aiwei Liu, Minghua He, Shaoxun Zeng, Sijun Zhang, Linhao Zhang, Chuhan Wu, Wei Jia, Yuan Liu, Zhou Xiao, Jie Zhou
Abstract
Autoregressive (AR) generation is the standard decoding paradigm for Large Language Models (LLMs), but its token-by-token nature limits parallelism at inference time. Diffusion Language Models (DLLMs) offer parallel decoding by recovering multiple masked tokens per step; however, in practice they often fail to translate this parallelism into deployment speed gains over optimized AR engines (e.g., vLLM). A key reason is that many DLLMs rely on bidirectional attention, which breaks standard prefix KV caching and forces repeated contextualization, undermining efficiency. We propose WeDLM, a diffusion decoding framework built entirely on standard causal attention to make parallel generation prefix-cache friendly. The core idea is to let each masked position condition on all currently observed tokens while keeping a strict causal mask, achieved by Topological Reordering that moves observed tokens to the physical prefix while preserving their logical positions. Building on this property, we introduce a streaming decoding procedure that continuously commits confident tokens into a growing left-to-right prefix and maintains a fixed parallel workload, avoiding the stop-and-wait behavior common in block diffusion methods. Experiments show that WeDLM preserves the quality of strong AR backbones while delivering substantial speedups, approaching 3× on challenging reasoning benchmarks and up to 10× in low-entropy generation regimes; critically, our comparisons are against AR baselines served by vLLM under matched deployment settings, demonstrating that diffusion-style decoding can outperform an optimized AR engine in practice.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext f84c6fa6-9c19-4b1e-9a5e-fe2f52d698bdCited by top-tier papers4
- d-TreeRPO: Towards More Reliable Policy Optimization for Diffusion Language ModelsLeyi Pan, Shuchang Tao, Yunpeng Zhai, Zheyu Fu et al.ACL 2026 · 14 citations
- DiffuReason: Enhancing Reasoning Ability for Diffusion Language Models via Monte Carlo Tree SearchYIPING SONG, Jinyu You, Zhiliang Tian, Jinsong Su et al.ICML 2026
- SPEED: Sharpened-Teacher Distillation for Parallel Decoding of Diffusion Language ModelsQiuhong Shen, Xingyi Yang, Xinyin Ma, Gongfan Fang et al.ICML 2026
- Lookahead-Then-Verify: Reliable Constrained Decoding for Diffusion LLMs under Context-Free GrammarsYitong Zhang, Yongmin Li, Yuetong Liu, Jia Li et al.ISSTA 2026
Builds on10
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code GenerationJiawei Liu, Chunqiu Steven Xia, Yuyao Wang, Lingming ZhangNeurIPS 2023 · 2,317 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
Related papers
- Trajectory-Level Speculative Decoding for Diffusion Language ModelsTianxiang Pan, Baitao Gong, Mo Guang, Hongwei Yong et al.ICML 2026
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 114 citations
- Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion ForcingXu Wang, Chenkai Xu, Yijie Jin, Jiachun Jin et al.ICLR 2026 · 116 citations
- d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory DistillationYu-Yang Qian, Junda Su, Lanxiang Hu, Peiyuan Zhang et al.ICML 2026 · 33 citations
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel DecodingChengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu et al.ICLR 2026 · 428 citations
