Parallel Sampling from Masked Diffusion Models via Conditional Independence Testing
Iskander Azangulov, Teodora Pandeva, Niranjani Prasad, Javier Zazo, Sushrut Karmalkar
摘要
Masked diffusion models (MDMs) offer a compelling alternative to autoregres- sive models (ARMs) for discrete text generation because they enable parallel token sampling, rather than sequential, left-to-right generation. This means po- tentially much faster inference. However, effective parallel sampling faces two competing requirements: (i) simultaneously updated tokens must be conditionally independent, and (ii) updates should prioritise high-confidence predictions. These goals conflict because high-confidence predictions often cluster and depend on each other, opportunities for parallel updates.
We present PUNT, a model-agnostic sampler that reconciles this trade-off. Our method identifies token dependencies and removes lower-confidence tokens from conflicting groups. This produces sets of indices for unmasking that satisfy both independence and confidence criteria. Our approach ensures improved parallel unmasking through approximate conditional independence testing.
Our experiments show that PUNT delivers a superior trade-off between accuracy and compute when compared to other strong training-free baselines, especially for generation of longer sequences. On the IFEval benchmark, it achieves up to 16% higher accuracy over baseline methods, including sequential generation (one-by- one). These gains hold across different values of hyperparameters, mitigating the need for brittle hyperparameter tuning. Moreover, we observe that PUNT induces an emergent hierarchical generation strategy, where the model first establishes high-level paragraph structure before local refinement, suggesting a planning-like generation process that contributes to strong alignment performance.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Learning Unmasking Policies for Diffusion Language ModelsMetod Jazbec, Theo X. Olausson, Louis Béthune, Pierre Ablin 等ICML 2026 · 被引用 24 次
- Tuning the Implicit Regularizer of Masked Diffusion Language Models: Enhancing Generalization via Insights from -ParityJianhao Huang, Baharan MirzasoleimanICML 2026 · 被引用 2 次
- Scheduling Thoughts: Learning the Order of Thought in Diffusion Language ModelsJiawei Xu, Minghui Liu, Aakriti Agrawal, Yifan Chen 等ICML 2026 · 被引用 1 次
- Unifying Masked Diffusion Models with Various Generation Orders and BeyondChunsan Hong, Sanghyun Lee, Jong Chul YEICML 2026
它引用的顶会 Paper19
- Structured Denoising Diffusion Models in Discrete State-SpacesJacob Austin, Daniel D. Johnson, Jonathan Ho, Daniel Tarlow 等NeurIPS 2021 · 被引用 2,256 次
- Fast Inference from Transformers via Speculative DecodingYaniv Leviathan, Matan Kalman, Yossi MatiasICML 2023 · 被引用 1,472 次
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang 等NeurIPS 2025 · 被引用 949 次
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan 等NeurIPS 2024 · 被引用 929 次
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 被引用 473 次
相关 Paper
- Partition Generative Modeling: Masked Modeling Without MasksJustin Deschenaux, Lan Tran, Caglar GulcehreICLR 2026 · 被引用 6 次
- Plan for Speed: Dilated Scheduling for Masked Diffusion Language ModelsOmer Luxembourg, Haim Permuter, Eliya NachmaniICML 2026 · 被引用 30 次
- Saber: Efficient Sampling with Adaptive Acceleration and Backtracking Enhanced Remasking for Diffusion Language Model in Code GenerationYihong Dong, Zhaoyu Ma, Xue Jiang, Zhiyuan Fan 等ACL 2026
- Masked Diffusion Models are Secretly Time-Agnostic Masked Models and Exploit Inaccurate Categorical SamplingKaiwen Zheng, Yongxin Chen, Hanzi Mao, Ming-Yu Liu 等ICLR 2025
- Improving Sampling for Masked Diffusion Models via Information GainKaisen Yang, Jayden Teoh, Kaicheng Yang, Yitong Zhang 等ICML 2026 · 被引用 6 次
