Efficient Perplexity Bound and Ratio Matching in Discrete Diffusion Language Models
Etrit Haxholli, Yeti Ziya Gurbuz, Ogul Can, Eli Waxman
摘要
While continuous diffusion models excel in modeling continuous distributions, their application to categorical data has been less effective. Recent work has shown that ratio-matching through score-entropy within a continuous-time discrete Markov chain (CTMC) framework serves as a competitive alternative to autoregressive models in language modeling. To enhance this framework, we first introduce three new theorems concerning the KL divergence between the data and learned distribution. Our results serve as the discrete counterpart to those established for continuous diffusion models and allow us to derive an improved upper bound of the perplexity. Second, we empirically show that ratio-matching performed by minimizing the denoising cross-entropy between the clean and corrupted data enables models to outperform those utilizing score-entropy with up to 10% lower perplexity/generative-perplexity, and 15% faster training steps. To further support our findings, we introduce and evaluate a novel CTMC transitionrate matrix that allows prediction refinement, and derive the analytic expression for its matrix exponential which facilitates the computation of conditional ratios thus enabling efficient training and generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Scaling Behavior of Discrete Diffusion Language ModelsDimitri von Rütte, Janis Fluri, Omead Pooladzandi, Bernhard Schölkopf 等ICLR 2026 · 被引用 34 次
- Partition Generative Modeling: Masked Modeling Without MasksJustin Deschenaux, Lan Tran, Caglar GulcehreICLR 2026 · 被引用 6 次
- Minibatch Optimal Transport and Perplexity Bound Estimation in Discrete Flow MatchingEtrit Haxholli, Yeti Z. Gurbuz, Oğul Can, Eli WaxmanICML 2026 · 被引用 3 次
- Don't Let It Fade: Preserving Edits in Diffusion Language Models via Token Timestep AllocationWoojin Kim, Jaeyoung DoNeurIPS 2025 · 被引用 2 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 被引用 5,234 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
相关 Paper
- Discrete Diffusion Modeling by Estimating the Ratios of the Data DistributionAaron Lou, Chenlin Meng, Stefano ErmonICML 2024 · 被引用 473 次
- Score-based Continuous-time Discrete Diffusion ModelsHaoran Sun, Lijun Yu, Bo Dai, Dale Schuurmans 等ICLR 2023 · 被引用 7 次
- Continuous Diffusion Model for Language ModelingJaehyeong Jo, Sung Ju HwangNeurIPS 2025 · 被引用 30 次
- Target Concrete Score Matching: A Holistic Framework for Discrete DiffusionRuixiang Zhang, Shuangfei Zhai, Yizhe Zhang, James Thornton 等ICML 2025
- A Continuous Time Framework for Discrete Denoising ModelsAndrew Campbell, Joe Benton, Valentin De Bortoli, Thomas Rainforth 等NeurIPS 2022 · 被引用 496 次
