FOCUS: DLLMs Know How to Tame Their Compute Bound
Kaihua Liang, Xin Tan, An Zhong, Hong Xu, Marco Canini
Abstract
Diffusion Large Language Models ( DLLMs ) offer a compelling alternative to Auto-Regressive models, but their deployment is constrained by high decoding cost. In this work, we identify a key inefficiency in DLLM decoding: while computation is parallelized over token blocks, only a small subset of tokens is decodable at each diffusion step, causing most compute to be wasted on non-decodable tokens. We further observe a strong correlation between attention-derived token importance and token-wise decoding probability. Based on this insight, we propose FOCUS, an inference system designed for DLLMs. By dynamically focusing computation on decodable tokens and evicting non-decodable ones on-the-fly, FOCUS increases the effective batch size, alleviating compute limitations and enabling scalable throughput. Empirical evaluations demonstrate that FOCUS achieves up to 3.52× throughput improvement over the production-grade engine LMDeploy in large-batch settings, while preserving or improving generation quality across multiple benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c7313a41-d88d-4788-83ed-18a0e4ba8174Builds on17
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 2,600 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han et al.ICLR 2024 · 1,714 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
Related papers
- Focus-dLLM: Accelerating Long-Context Diffusion LLM Inference via Confidence-Guided Context FocusingLingkun Long, Yushi Huang, Shihao Bai, Ruihao Gong et al.ACL 2026 · 3 citations
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 114 citations
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast InferenceAiwei Liu, Minghua He, Shaoxun Zeng, Sijun Zhang et al.ICML 2026 · 37 citations
- ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-SkippingZijian Zhu, Fei Ren, Zhanhong Tan, Kaisheng MaICLR 2026 · 7 citations
- DyLLM: Efficient Diffusion LLM Inference via Saliency-based Token Selection and Partial AttentionYounjoo Lee, Seungkyun Dan, Junghoo Lee, Jaiyoung Park et al.ICML 2026 · 2 citations
