UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion Models
Jiaqi Wang, Haoge Deng, Ting Pan, Yang Liu, Chengyuan Wang, Fan Zhang, Yonggang Qi, Xinlong Wang
摘要
Uniform Discrete Diffusion (UDM) has recently emerged as a promising paradigm for discrete generative modeling; however, its integration with reinforcement learning remains largely unexplored. We observe that naively adapting GRPO to UDM leads to unstable training and marginal performance. To address this, we propose , the first framework that integrates UDM with RL. Our method is guided by two key insights: (i) treating the final clean sample, rather than intermediate predicted sample, as the action provides more accurate and stable optimization signals; and (ii) adopting the forward process to reconstruct the training trajectories helps the model learn probability paths that are more consistent with pretraining. For efficiency, we introduce Reduction-Step and CFG-Free training strategies. significantly improves the performance of the base model across multiple T2I tasks. Notably, GenEval accuracy improves from to and PickScore increases from to , achieving state-of-the-art performance in both continuous and discrete settings. On the OCR benchmark, accuracy improves from to , further validating the effectiveness and generalization capability of our method.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper23
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
相关 Paper
- Graph-GRPO: Training Graph Flow Models with Reinforcement LearningBaoheng Zhu, Deyu Bo, Delvin Zhang, Xiao WangICML 2026 · 被引用 3 次
- Fine-Grained GRPO for Precise Preference Alignment in Flow ModelsYujie Zhou, Pengyang Ling, Jiazi Bu, Yibin Wang 等CVPR 2026 · 被引用 19 次
- Consolidating Reinforcement Learning for Multimodal Discrete Diffusion ModelsTianren Ma, Mu Zhang, Yibing Wang, Qixiang YeICLR 2026 · 被引用 10 次
- RebRL: Reinforcing Discrete Visual Diffusion Models with Rebalanced Timestep CreditsMu Zhang, Tianren Ma, Yunfan Liu, Kun Hu 等CVPR 2026
- Unified Multimodal Models as Auto-EncodersZhiyuan Yan, Kaiqing Lin, Zongjian Li, Junyan Ye 等CVPR 2026 · 被引用 12 次
