DecoCal: Decoding with Calibration in Diffusion Large Language Models
Fan Xu, Huixuan Zhang, Xiaojun Wan
摘要
Diffusion Large Language Models (DLLMs) generate text via iterative masked-token denoising, supporting parallel prediction and bidirectional context modeling. Despite these advantages, decoding remains challenging: many tokens appear predictable early, yet single-step predictions are often unstable, exhibiting temporal oscillations or overconfidence, making it difficult to determine which tokens can be safely committed. To address these challenges, we propose DecoCal 1 , a Decoding framework that explicitly performs Calibration of tokenlevel confidence across diffusion steps and leverages the calibrated results to guide decoding decisions. Specifically, DecoCal aggregates historical predictions to maintain calibrated confidence, triggering unmasking only when a token is sufficiently stable, while a remasking mechanism allows revision of premature commitments. This calibration-based design enables early decoding of reliably converged tokens while deferring or correcting unstable ones, balancing reliability and speed. Experiments on multiple DLLMs and benchmarks show that DecoCal improves generation accuracy compared to existing strategies. Our results highlight the importance of temporal calibration in unlocking the full potential of diffusion-based language generation. 𝐩 ! (#) 𝐪 !%& (#) 𝐪 ! (#) 𝑥 (') [MASK] 𝑥 (() [MASK] 𝐾𝐿 a. Confidence Check b. Consistency Check Calibration Sequence [MASK] 𝑥 (') 𝑥 (&) 𝑥 (() 𝑥 ()) [MASK] Unmask Remask 𝑥 (') [MASK] 𝑥 (() [MASK] [MASK] 𝑥 (') 𝑥 (&) [MASK] 𝑥 ()) [MASK] Cool-down Delayed Consistency Check
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper11
- Diffusion-LM Improves Controllable Text GenerationXiang Lisa Li, John Thickstun, Ishaan Gulrajani, Percy Liang 等NeurIPS 2022 · 被引用 1,546 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Self-Consistency Improves Chain of Thought Reasoning in Language ModelsXuezhi Wang, Jason Wei, Dale Schuurmans, Quoc V. Le 等ICLR 2023 · 被引用 681 次
- MaskGIT: Masked Generative Image TransformerHuiwen Chang, Han Zhang, Lu Jiang, Ce Liu 等CVPR 2022 · 被引用 346 次
- Latent Diffusion for Language GenerationJustin Lovelace, Varsha Kishore, Chao Wan, Eliot Shekhtman 等NeurIPS 2023 · 被引用 177 次
相关 Paper
- CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace CreditKangyu Wang, Zhiyun Jiang, Haibo Feng, Weijia Zhao 等ACL 2026 · 被引用 11 次
- Don't Settle Too Early: Self-Reflective Remasking for Diffusion Language ModelsZemin Huang, Yuhang Wang, Zhiyang Chen, Guo-Jun QiICLR 2026 · 被引用 40 次
- Efficient Diffusion LLMs via Temporal-Spatial Parallel Decoding and Confidence ExtrapolationZekai Li, Ji Liu, Yiqing Huang, Ziqiong Liu 等ICML 2026 · 被引用 1 次
- Hierarchy Decoding: A Training-free Parallel Decoding Strategy for Diffusion Large Language ModelsXiaojing Qi, Lun Du, Xinyuan Zhang, Lanning Wei 等ICLR 2026
- FAIR-Calib: Frontier-Aware Instability-Reweighted Calibration for Post-Training Quantization of Diffusion Large Language ModelsHaoyu Huang, Linlin Yang, Sheng Xu, Boyu Liu 等ICML 2026
