Lune

ICLR2026Top-tier venue

Hierarchy Decoding: A Training-free Parallel Decoding Strategy for Diffusion Large Language Models

Xiaojing Qi, Lun Du, Xinyuan Zhang, Lanning Wei, Tao Jin, Da Zheng

2026Year
1Top-tier citations

Abstract

The utilization of large language models (LLMs) has become increasingly widespread, and has attracted considerable attention. Although Discrete Diffusion Large Language Models (dLLMs) mitigate the inference latency inherent in autoregressive LLM decoding, their computational overhead remains substantial. To address this challenge, we propose Hierarchy-dLLM, a hierarchical decoding framework inspired by the divide-and-conquer principle. Our method recursively partitions masked spans into smaller sub-decoding areas and decodes tokens according to their confidence, which substantially increases the number of tokens generated per forward pass and improves information utilization. Extensive experiments conducted on multiple benchmarks demonstrate that Hierarchy-dLLM achieves accuracy comparable to or even surpassing existing baselines. Meanwhile, it is up to 17× faster than vanilla decoding and about 1.5× faster than the Fast-dLLM. These results establish hierarchical decoding as a practical solution for efficient dLLMs inference. The implementation is available at https://github.com/inclusionAI/dInfer .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext 5115bda0-20b0-4e7c-bfa7-dc5e62fe7004

Cited by top-tier papers1

Ask how each one uses it

Builds on14

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines