Lune

ICLR2026顶会

DPad: Efficient Diffusion Language Models with Suffix Dropout

Xinhua Chen, Sitao Huang, Cong Guo, Chiyue Wei, Yintao He, Jianyi Zhang, Hai Li, Yiran Chen

2026年份
44被引次数
8顶会引用

摘要

Diffusion-based Large Language Models (dLLMs) parallelize text generation by framing decoding as a denoising process, but suffer from high computational overhead since they predict all future suffix tokens at each step while retaining only a small fraction. We propose Diffusion Scratchpad(DPad)\textbf{Diffusion Scratchpad} (\textbf{\textit{DPad}}), a training-free method that restricts attention to a structured subset of suffix tokens, preserving fidelity while eliminating redundancy. DPad\textit{DPad} integrates two strategies: (i) a sliding window\textit{sliding window}, which maintains a fixed-length suffix window, and (ii) distance-decay dropout\textit{distance-decay dropout}, which deterministically removes distant suffix tokens before attention computation. This concise design is compatible with existing optimizations such as parallel decoding and prefix caching, and lends itself to a lightweight implementation. Comprehensive evaluations across multiple benchmarks on LLaDA\texttt{LLaDA} and Dream\texttt{Dream} models demonstrate that DPad\textit{DPad} delivers up to 61.4×\mathbf{61.4\times} speedup over vanilla dLLMs while maintaining comparable accuracy, highlighting its potential for efficient and scalable long-sequence inference.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper8

问问它们各自怎么用它

它引用的顶会 Paper20

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖