Towards Efficient and Effective Diffusion Language Model Inference via Semantic-Aware Adaptive Denoising
Fan Li, Yu Gu, Zhigang Wang, Fangling Leng, Zhenghao Liu, Ge Yu
摘要
Diffusion language models (DLMs) have emerged as a powerful non-autoregressive alternative to GPT-style sequential generation, but suffer from substantial computational overhead due to their iterative parallel denoising. Existing acceleration works cannot accurately detect semantically stabilized tokens and then skip computation, leading to sub-optimal speedup in practice. This paper presents the first systematic study of convergence dynamics in DLMs. Innovative observations include the misalignment between traditionally used scalar detection criterion and the semantic convergence, and the post-peak confidence score, that wastes denoising computation and degrades inference quality. To address these limitations, we propose Ada-DLM, a semantic-aware adaptive denoising framework that encodes the trajectory of scalar confidence scores into an evolutionaware feature vector and then clusters vectors proactively to adaptively identify semantically converged tokens. Furthermore, we incorporate system-level optimizations to maximize runtime efficiency. Experiments show that Ada-DLM consistently outperforms the SOTA competitor, achieving up to 2× speedup and 19% quality improvement. That offers a practical path toward efficient high-quality DLM deployment.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
相关 Paper
- Dynamic-dLLM: Dynamic Cache-Budget and Adaptive Parallel Decoding for Training-Free Acceleration of Diffusion LLMTianyi Wu, Xiaoxi Sun, Yanhua Jiao, Yulin Li 等ICLR 2026 · 被引用 6 次
- ES-dLLM: Efficient Inference for Diffusion Large Language Models by Early-SkippingZijian Zhu, Fei Ren, Zhanhong Tan, Kaisheng MaICLR 2026 · 被引用 7 次
- AdaBlock-dLLM: Semantic-Aware Diffusion LLM Inference via Adaptive Block SizeGuanxi Lu, Hao Mark Chen, Yuto Karashima, Zhican Wang 等ICLR 2026 · 被引用 32 次
- WavefrontDiffusion: Dynamic Decoding Schedule for Improved ReasoningHaojin Yang, Rui Hu, Zequn Sun, Rui Zhou 等ICLR 2026 · 被引用 5 次
- Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence ExtrapolationZekai Li, Ji Liu, Yiqing Huang, Ziqiong Liu 等ICML 2026 · 被引用 1 次
