Lune

ICLR2026顶会

Hierarchy Decoding: A Training-free Parallel Decoding Strategy for Diffusion Large Language Models

Xiaojing Qi, Lun Du, Xinyuan Zhang, Lanning Wei, Tao Jin, Da Zheng

出版方
2026年份
1顶会引用

摘要

The utilization of large language models (LLMs) has become increasingly widespread, and has attracted considerable attention. Although Discrete Diffusion Large Language Models (dLLMs) mitigate the inference latency inherent in autoregressive LLM decoding, their computational overhead remains substantial. To address this challenge, we propose Hierarchy-dLLM, a hierarchical decoding framework inspired by the divide-and-conquer principle. Our method recursively partitions masked spans into smaller sub-decoding areas and decodes tokens according to their confidence, which substantially increases the number of tokens generated per forward pass and improves information utilization. Extensive experiments conducted on multiple benchmarks demonstrate that Hierarchy-dLLM achieves accuracy comparable to or even surpassing existing baselines. Meanwhile, it is up to 17× faster than vanilla decoding and about 1.5× faster than the Fast-dLLM. These results establish hierarchical decoding as a practical solution for efficient dLLMs inference. The implementation is available at https://github.com/inclusionAI/dInfer .

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper14

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖