Mosaic: Unlocking Over 30 Context Length for Diffusion LLMs Inference via Global Memory Planning and Dynamic Peak Taming
Liang Zheng, Bowen Shi, Yitao Hu, Jiawei Zhang, Ruofan Li, Guotao Yang, Zhixin Zhao, Zhengchao Wang, Sheng Chen, Wenxin Li, Dezhi Ran, Tao Xie, Keqiu Li
摘要
Diffusion-based large language models (dLLMs) have emerged as a promising alternative to autoregressive models, leveraging simultaneous denoising to enable global planning and iterative refinement. These properties make dLLMs attractive for long-context generation. However, deploying dLLMs faces a prohibitive memory barrier, as existing inference systems are inefficient for the diffusion paradigm. We observe that current inference systems are misaligned with dLLMs. Unlike autoregressive models, whose memory footprint is dominated by the KV-Cache, dLLMs are bottlenecked by transient activations rematerialized per step. Moreover, generic memory reuse mechanisms lack the global visibility to handle dynamic memory peaks of dLLMs, which alternate between logits and feed-forward networks. To address these challenges, we present Mosaic, a memory-efficient inference system that shifts dLLM execution from local, static management to a global, dynamic paradigm. Mosaic integrates (i) a mask-only logits kernel eliminating redundant activation materialization, (ii) a lazy chunking optimizer using online heuristics to tame dynamic memory peaks, and (iii) a global memory manager leveraging virtual addressing to mitigate memory fragmentation. Evaluations show that Mosaic reduces the memory peak-to-average ratio by 2.71 on average and increases the maximum inference sequence length on identical hardware by 15.30--32.34. Crucially, Mosaic is training-free and preserves exact model outputs, while reducing end-to-end latency by 2.5%--55.4%. Our code is publicly available at https://github.com/flashserve/Mosaic.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
- Efficient Memory Management for Large Language Model Serving with PagedAttentionWoosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng 等SOSP 2023 · 被引用 1,016 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
相关 Paper
- DiffuMamba: High-Throughput Diffusion LMs with Mamba BackboneVaibhav Singh, Oleksiy Ostapenko, Pierre-André Noël, Eugene Belilovsky 等ICML 2026
- dLLM-Cache: Accelerating Diffusion Large Language Models with Adaptive CachingZhiyuan Liu, Yicun Yang, Yaojie Zhang, Junjie Chen 等ICML 2026 · 被引用 156 次
- Beyond Fixed: Training-Free Variable-Length Denoising for Diffusion Large Language ModelsJinsong Li, Xiaoyi Dong, Yuhang Zang, Yuhang Cao 等ICLR 2026 · 被引用 34 次
- LoSA: Locality Aware Sparse Attention in Diffusion Language ModelsHaocheng Xi, Harman Singh, Yuezhou Hu, Coleman Hooper 等ICML 2026 · 被引用 2 次
- FlashDLM: Accelerating Diffusion Language Model Inference via Efficient KV Caching and Guided DiffusionZhanqiu Hu, Jian Meng, Yash Akhauri, Mohamed S. Abdelfattah 等ICLR 2026 · 被引用 56 次
