Lune

NeurIPS2025顶会

ReDi: Rectified Discrete Flow

Jaehoon Yoo, Wonjung Kim, Seunghoon Hong

2025年份
13被引次数
1顶会引用

摘要

Discrete Flow-based Models (DFMs) are powerful generative models for highquality discrete data but typically suffer from slow sampling speeds due to their reliance on iterative decoding processes. This reliance on a multi-step process originates from the factorization approximation of DFMs, which is necessary for handling high-dimensional data. In this paper, we analyze the factorization approximation error using Conditional Total Correlation (TC), and reveal its dependence on the coupling. To address the challenge of efficient few-step generation, we propose Rectified Discrete Flow (ReDi), a novel iterative method that reduces the underlying factorization error (measured as Conditional TC) by rectifying the coupling between source and target distributions. We theoretically prove that each ReDi step guarantees a monotonic decreasing Conditional TC, ensuring its convergence. Empirically, ReDi significantly reduces Conditional TC and enables few-step generation. Moreover, we demonstrate that the rectified couplings are well-suited for training efficient one-step models on image generation. ReDi offers a simple and theoretically grounded approach for tackling the few-step challenge, providing a new perspective on efficient discrete data synthesis. Code is available at https://github.com/Ugness/ReDi_discrete.

reveal its dependence on the coupling. Inspired by Rectified Flows [24,25,42] in continuous domain, we propose Rectified Discrete Flow (ReDi) to enable efficient few-step generation by rectifying the coupling of discrete data, which in turn reduces the Conditional TC. By focusing on coupling rectification, our method provides a simpler alternative to prior works [10,17,31], as it requires neither a specialized training strategy nor the handling of separate teacher-student models, which in turn reduces memory requirements. This simplicity enables ReDi to be broadly applicable to various DFMs including other distillation frameworks.

We demonstrated our method's effectiveness both theoretically and empirically. We theoretically prove that each ReDi iteration guarantees monotonically decreasing Conditional TC and empirically show that each rectification significantly reduces it. We evaluated our method on class conditional image generation and text generation. On image generation, ReDi shows comparable few-step generation performance against existing distillation methods, and significantly outperforms in onestep generation, due to direct rectification of the couplings contributing to the factorization error. On text generation, we observe that iteratively applying rectification improves the sampling efficiency and that ReDi can also be applied with existing distillation methods.

2 Related Works

Discrete flow-based models (DFMs) are used to generate discrete data such as images [2,7,8,16,39], videos [43,44], text [1, 29-31, 34, 35], and protein [6]. DFMs generate data by learning the flow from initial states (often set as masked states or uniform random states). The generative flow is learned by two primary formalisms, reversing a corruption process (e.g., masked generative models [2, 7, 8, 39], discrete diffusion [1,26,30,34,35]), or constructing bridges between initial distribution and data distribution (e.g., Discrete Flow Matching [6, 13], Schrödinger Bridges [21,22]). Although they show powerful performance on discrete data synthesis, and efficient sampling cost compared to autoregressive models as they support generating multiple states simultaneously [10,30,31,34], they still require slow multi-step decoding process for successive generation [17].

To address the slow sampling speeds of multi-step DFMs, prior works [10,17,31] have explored methods for distillation and faster generation. These approaches typically aim to distill a slower, multi-step teacher model into a faster, few-step student model. They primarily focus on modifying the training objective or designing specific training procedures tailored for distillation, and are sometimes specific to a particular DFM framework. For instance, SDTT [10] is tailored for masked diffusion models, and the dual consistency distillation method suggested in DUO [31] is tailored for uniform diffusion. While Di4C [17] suggested an objective function that is applicable for various discrete diffusion models, it utilizes four loss terms, requiring tuning of weights.

Alternative approaches [23,41] introduce auxiliary models to reduce decoding steps. Discrete Copula Diffusion [23], for instance, requires a pretrained autoregressive model as an additional copula model. EDLM [41] takes another approach, using an energy-based model to guide sampling; however, its practical sampling efficiency remains limited as its algorithm requires sampling multiple candidates and selecting the most probable one. In contrast to the prior works, our method improves few-step generation in DFMs by focusing on the coupling itself, rather than solely on modifying the training process

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper1

问问它们各自怎么用它

它引用的顶会 Paper29

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖