ReDi: Rectified Discrete Flow
Jaehoon Yoo, Wonjung Kim, Seunghoon Hong
Abstract
Discrete Flow-based Models (DFMs) are powerful generative models for highquality discrete data but typically suffer from slow sampling speeds due to their reliance on iterative decoding processes. This reliance on a multi-step process originates from the factorization approximation of DFMs, which is necessary for handling high-dimensional data. In this paper, we analyze the factorization approximation error using Conditional Total Correlation (TC), and reveal its dependence on the coupling. To address the challenge of efficient few-step generation, we propose Rectified Discrete Flow (ReDi), a novel iterative method that reduces the underlying factorization error (measured as Conditional TC) by rectifying the coupling between source and target distributions. We theoretically prove that each ReDi step guarantees a monotonic decreasing Conditional TC, ensuring its convergence. Empirically, ReDi significantly reduces Conditional TC and enables few-step generation. Moreover, we demonstrate that the rectified couplings are well-suited for training efficient one-step models on image generation. ReDi offers a simple and theoretically grounded approach for tackling the few-step challenge, providing a new perspective on efficient discrete data synthesis. Code is available at https://github.com/Ugness/ReDi_discrete.
reveal its dependence on the coupling. Inspired by Rectified Flows [24,25,42] in continuous domain, we propose Rectified Discrete Flow (ReDi) to enable efficient few-step generation by rectifying the coupling of discrete data, which in turn reduces the Conditional TC. By focusing on coupling rectification, our method provides a simpler alternative to prior works [10,17,31], as it requires neither a specialized training strategy nor the handling of separate teacher-student models, which in turn reduces memory requirements. This simplicity enables ReDi to be broadly applicable to various DFMs including other distillation frameworks.
We demonstrated our method's effectiveness both theoretically and empirically. We theoretically prove that each ReDi iteration guarantees monotonically decreasing Conditional TC and empirically show that each rectification significantly reduces it. We evaluated our method on class conditional image generation and text generation. On image generation, ReDi shows comparable few-step generation performance against existing distillation methods, and significantly outperforms in onestep generation, due to direct rectification of the couplings contributing to the factorization error. On text generation, we observe that iteratively applying rectification improves the sampling efficiency and that ReDi can also be applied with existing distillation methods.
2 Related Works
Discrete flow-based models (DFMs) are used to generate discrete data such as images [2,7,8,16,39], videos [43,44], text [1, 29-31, 34, 35], and protein [6]. DFMs generate data by learning the flow from initial states (often set as masked states or uniform random states). The generative flow is learned by two primary formalisms, reversing a corruption process (e.g., masked generative models [2, 7, 8, 39], discrete diffusion [1,26,30,34,35]), or constructing bridges between initial distribution and data distribution (e.g., Discrete Flow Matching [6, 13], Schrödinger Bridges [21,22]). Although they show powerful performance on discrete data synthesis, and efficient sampling cost compared to autoregressive models as they support generating multiple states simultaneously [10,30,31,34], they still require slow multi-step decoding process for successive generation [17].
To address the slow sampling speeds of multi-step DFMs, prior works [10,17,31] have explored methods for distillation and faster generation. These approaches typically aim to distill a slower, multi-step teacher model into a faster, few-step student model. They primarily focus on modifying the training objective or designing specific training procedures tailored for distillation, and are sometimes specific to a particular DFM framework. For instance, SDTT [10] is tailored for masked diffusion models, and the dual consistency distillation method suggested in DUO [31] is tailored for uniform diffusion. While Di4C [17] suggested an objective function that is applicable for various discrete diffusion models, it utilizes four loss terms, requiring tuning of weights.
Alternative approaches [23,41] introduce auxiliary models to reduce decoding steps. Discrete Copula Diffusion [23], for instance, requires a pretrained autoregressive model as an additional copula model. EDLM [41] takes another approach, using an energy-based model to guide sampling; however, its practical sampling efficiency remains limited as its algorithm requires sampling multiple candidates and selecting the most probable one. In contrast to the prior works, our method improves few-step generation in DFMs by focusing on the coupling itself, rather than solely on modifying the training process
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67d3bc78-49e1-4e0e-b16e-67ad3d335478Cited by top-tier papers1
Ask how each one uses itBuilds on29
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
- Large Language Diffusion ModelsShen Nie, Fengqi Zhu, Zebin You, Xiaolu Zhang et al.NeurIPS 2025 · 949 citations
- Simple and Effective Masked Diffusion Language ModelsSubham S. Sahoo, Marianne Arriola, Yair Schiff, Aaron Gokaslan et al.NeurIPS 2024 · 929 citations
- Muse: Text-To-Image Generation via Masked Generative TransformersHuiwen Chang, Han Zhang, Jarred Barber, Aaron Maschinot et al.ICML 2023 · 751 citations
Related papers
- InstaFlow: One Step is Enough for High-Quality Diffusion-Based Text-to-Image GenerationXingchao Liu, Xiwen Zhang, Jianzhu Ma, Jian Peng et al.ICLR 2024 · 358 citations
- Universal Inverse Distillation for Matching Models with Real-Data Supervision (No GANs)Nikita Kornilov, David Li, Tikhon Mavrin, Aleksei Leonov et al.ICLR 2026 · 5 citations
- Corrected Samplers for Discrete Flow ModelsZhengyan Wan, Yidong Ouyang, Liyan Xie, Hongyuan Zha et al.ICML 2026
- Improving the Training of Rectified FlowsSangyun Lee, Zinan Lin, Giulia FantiNeurIPS 2024 · 119 citations
- Adaptive Piecewise Distillation for Efficient LiDAR Data GenerationRuibo Li, Xiaofeng Yang, Ze Yang, Jiacheng Wei et al.AAAI 2026
