From Bits to Rounds: Parallel Decoding with Exploration for Diffusion Language Models
Hengyu Fu, Baihe Huang, Virginia Adams, Charles Wang, Junkeun Yi, Mohammad Mahdi Kamani, Venkat Krishna Srinivasan, Jiantao Jiao
Abstract
Diffusion Language Models (DLMs) have recently emerged as a strong alternative to autoregressive language models (LMs). DLMs offer comparable accuracy with faster inference speed via parallel decoding. However, standard DLM decoding strategies relying on high-confidence tokens encounter an inherent information-theoretic bottleneck that restricts decoding progress and ultimately slows generation. We demonstrate both theoretically and empirically that prioritizing high-confidence tokens is inherently inefficient. High-probability tokens carry negligible information and strictly relying on them limits the effective progress made in each decoding round. We prove that the number of decoding rounds must grow linearly with the sample's total information (negative log-likelihood) and inversely with the per-round information budget, establishing a bits-to-rounds principle. We also propose Explore-Then-Exploit (ETE), a training-free decoding strategy that maximizes information throughput and decoding efficiency. ETE combines cross-block decoding with targeted exploration of high-uncertainty tokens to reshape the conditional distribution and trigger cascades of confident predictions. Experiments verify our theoretical bounds and demonstrate that ETE consistently reduces the required number of decoding rounds compared to confidence-only baselines without compromising generation quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Can I Have Your Order? Monte-Carlo Tree Search for Slot Filling Ordering in Diffusion Language ModelsJoshua Ong, Yu Zhao, Mihaela C. Stoian, Wenda Li et al.ICML 2026
- Residual Context Diffusion Language ModelsYuezhou Hu, Harman Singh, Monishwaran Maheswaran, Haocheng Xi et al.ICML 2026
Builds on21
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow et al.NeurIPS 2021 · 2,256 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz et al.NeurIPS 2024 · 751 citations
Related papers
- Trajectory-Level Speculative Decoding for Diffusion Language ModelsTianxiang Pan, Baitao Gong, Mo Guang, Hongwei Yong et al.ICML 2026
- Search or Accelerate: Confidence-Switched Position Beam Search for Diffusion Language ModelsMingyu Cao, Alvaro Correia, Christos Louizos, Shiwei Liu et al.ICML 2026
- Accelerating Diffusion LLMs via Adaptive Parallel DecodingDaniel Israel, Guy Van den Broeck, Aditya GroverNeurIPS 2025 · 114 citations
- WeDLM: Reconciling Diffusion Language Models with Standard Causal Attention for Fast InferenceAiwei Liu, Minghua He, Shaoxun Zeng, Sijun Zhang et al.ICML 2026 · 37 citations
- CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace CreditKangyu Wang, Zhiyun Jiang, Haibo Feng, Weijia Zhao et al.ACL 2026 · 11 citations
