FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided Diffusion
Zhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah, Jae-sun Seo, Zhiru Zhang, Udit Gupta
Abstract
Diffusion language models offer parallel token generation and inherent bidirectionality, promising more efficient and powerful sequence modeling compared to autoregressive approaches. However, state-of-the-art diffusion models (e.g., Dream 7B, LLaDA 8B) suffer from slow inference. While they match the quality of similarly sized Autoregressive (AR) Models (e.g., Qwen2.5 7B, Llama3 8B), their iterative denoising requires multiple full-sequence forward passes, resulting in high computational costs and latency, particularly for long input prompts and long-context scenarios. Furthermore, parallel token generation introduces token incoherence problems, and current sampling heuristics suffer from significant quality drops with decreasing denoising steps. We address these limitations with two training-free techniques. First, we propose FreeCache, a Key-Value (KV) approximation caching technique that reuses stable KV projections across denoising steps, effectively reducing the computational cost of DLM inference. Second, we introduce Guided Diffusion, a training-free method that uses a lightweight pretrained autoregressive model to supervise token unmasking, dramatically reducing the total number of denoising iterations without sacrificing quality. We conduct extensive evaluations on open-source reasoning benchmarks, and our combined methods deliver an average of 12.14 end-to-end speedup across various tasks with negligible accuracy degradation. For the first time, diffusion language models achieve a comparable and even faster latency as the widely adopted autoregressive models. Our work successfully paved the way for scaling up the diffusion language model to a broader scope of applications across different domains. Our code and implementation are available at https://github.com/ZhanqiuHu/flash-dlm-experimental.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7ff3b8e7-0116-46bb-8632-721514e7f9f7Cited by top-tier papers4
- Plan for Speed: Dilated Scheduling for Masked Diffusion Language ModelsOmer Luxembourg, Haim Permuter, Eliya NachmaniICML 2026 · 30 citations
- Planned DiffusionDaniel Mingyi Israel, Tian Jin, Ellie Y Cheng, Guy Van den Broeck et al.ICLR 2026 · 8 citations
- Locally Coherent Parallel Decoding in Diffusion Language ModelsMichael Hersche, Nicolas Menet, Ronan Tanios, Abbas RahimiICML 2026 · 1 citation
- Lookahead-Then-Verify: Reliable Constrained Decoding for Diffusion LLMs under Context-Free GrammarsYitong Zhang, Yongmin Li, Yuetong Liu, Jia Li et al.ISSTA 2026
Builds on13
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 1,472 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
Related papers
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingZhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen et al.ICML 2026 · 156 citations
- dCache: Accelerating Diffusion-Based LLMs via Dual Adaptive CachingYuchu Jiang, Yue Cai, Xiangzhong Luo, Jiale Fu et al.ICLR 2026 · 16 citations
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li et al.ICLR 2026 · 6 citations
- dKV-Cache: The Cache for Diffusion Language ModelsXinyin Ma, Runpeng Yu, Gongfan Fang, Xinchao WangNeurIPS 2025 · 145 citations
- Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel DecodingChengyue Wu, Hao Zhang, Shuchen Xue, Zhijian Liu et al.ICLR 2026 · 428 citations
