Sparse-dLLM: Accelerating Diffusion LLMs with Dynamic Cache Eviction
Yuerong Song, Xiaoran Liu, Ruixiao Li, Zhigeng Liu, Zengfeng Huang, Qipeng Guo, Ziwei He, Xipeng Qiu
Abstract
Diffusion Large Language Models (dLLMs) enable breakthroughs in reasoning and parallel decoding but suffer from prohibitive quadratic computational complexity and memory overhead during inference. Current caching techniques accelerate decoding by storing full-layer states, yet impose substantial memory usage that limit long-context applications. Our analysis of attention patterns in dLLMs reveals persistent cross-layer sparsity, with pivotal tokens remaining salient across decoding steps and low-relevance tokens staying unimportant, motivating selective cache eviction. We propose Sparse-dLLM, the first training-free framework integrating dynamic cache eviction with sparse attention via delayed bidirectional sparse caching. By leveraging the stability of token saliency over steps, it retains critical tokens and dynamically evicts unimportant prefix/suffix entries using an attention-guided strategy. Extensive experiments on LLaDA and Dream series demonstrate Sparse-dLLM achieves up to 10× higher throughput than vanilla dLLMs, with comparable performance and similar peak memory costs, outperforming previous methods in efficiency and effectiveness. The code is available at https://github.com/OpenMOSS/Sparse-dLLM .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 480d2ae9-6c0e-4aa0-a33f-fd92e9e56bd3Cited by top-tier papers10
- Accelerating Diffusion Large Language Models with SlowFast Sampling: The Three Golden PrinciplesQingyan Wei, Yaojie Zhang, Zhiyuan Liu, Puyu Zeng et al.ICLR 2026 · 44 citations
- TEAM: Temporal–Spatial Consistency Guided Expert Activation for MoE Diffusion Language Model AccelerationLINYE WEI, Zixiang Luo, Pingzhi Tang, Meng LiICML 2026 · 7 citations
- Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language ModelsShufan Li, Jiuxiang Gu, Kangning Liu, Zhe Lin et al.CVPR 2026 · 6 citations
- Stop the Flip-Flop: Context-Preserving Verification for Fast Revocable Diffusion DecodingYanzheng Xiang, Lan Wei, Yizhen Yao, Qinglin Zhu et al.ICML 2026 · 3 citations
- LoSA: Locality Aware Sparse Attention in Diffusion Language ModelsHaocheng Xi, Harman Singh, Yuezhou Hu, Coleman Hooper et al.ICML 2026 · 2 citations
Builds on13
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou et al.ICLR 2021 · 7,905 citations
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao et al.AAAI 2020 · 2,916 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- SnapKV: LLM Knows What You are Looking for Before GenerationYuhong Li, Yingbing Huang, Bowen Yang, Bharat Venkitesh et al.NeurIPS 2024 · 1,019 citations
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen et al.NeurIPS 2023 · 1,003 citations
Related papers
- DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial AttentionYounjoo Lee, Seungkyun Dan, Junghoo Lee, Jaiyoung Park et al.ICML 2026 · 2 citations
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li et al.ICLR 2026 · 6 citations
- ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-SkippingZijian Zhu, Fei Ren, Zhanhong Tan, Kaisheng MaICLR 2026 · 7 citations
- dCache: Accelerating Diffusion-Based LLMs via Dual Adaptive CachingYuchu Jiang, Yue Cai, Xiangzhong Luo, Jiale Fu et al.ICLR 2026 · 16 citations
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingZhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen et al.ICML 2026 · 156 citations
