Mosaic: Unlocking Over 30 Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming
Liang Zheng, Bowen Shi, Yitao Hu, Jiawei Zhang, Ruofan Li, Guotao Yang, Zhixin Zhao, Zhengchao Wang, Sheng Chen, Wenxin Li, Dezhi Ran, Tao Xie, Keqiu Li
Abstract
Diffusion-based large language models (dLLMs) have emerged as a promising alternative to autoregressive models, leveraging simultaneous denoising to enable global planning and iterative refinement. These properties make dLLMs attractive for long-context generation. However, deploying dLLMs faces a prohibitive memory barrier, as existing inference systems are inefficient for the diffusion paradigm. We observe that current inference systems are misaligned with dLLMs. Unlike autoregressive models, whose memory footprint is dominated by the KV-Cache, dLLMs are bottlenecked by transient activations rematerialized per step. Moreover, generic memory reuse mechanisms lack the global visibility to handle dynamic memory peaks of dLLMs, which alternate between logits and feed-forward networks. To address these challenges, we present Mosaic, a memory-efficient inference system that shifts dLLM execution from local, static management to a global, dynamic paradigm. Mosaic integrates (i) a mask-only logits kernel eliminating redundant activation materialization, (ii) a lazy chunking optimizer using online heuristics to tame dynamic memory peaks, and (iii) a global memory manager leveraging virtual addressing to mitigate memory fragmentation. Evaluations show that Mosaic reduces the memory peak-to-average ratio by 2.71 on average and increases the maximum inference sequence length on identical hardware by 15.30--32.34. Crucially, Mosaic is training-free and preserves exact model outputs, while reducing end-to-end latency by 2.5%--55.4%. Our code is publicly available at https://github.com/flashserve/Mosaic.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 75958b5a-5f34-4bf9-959e-bc4758f2d331Builds on13
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang et al.NeurIPS 2022 · 1,546 citations
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng et al.SOSP 2023 · 1,016 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
Related papers
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky et al.ICML 2026
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingZhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen et al.ICML 2026 · 156 citations
- Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language ModelsJinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao et al.ICLR 2026 · 34 citations
- LoSA: Locality Aware Sparse Attention in Diffusion Language ModelsHaocheng Xi, Harman Singh, Yuezhou Hu, Coleman Hooper et al.ICML 2026 · 2 citations
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided DiffusionZhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah et al.ICLR 2026 · 56 citations
