RebRL: Reinforcing Discrete Visual Diffusion Models with Rebalanced Timestep Credits
Mu Zhang, Tianren Ma, Yunfan Liu, Kun Hu, Qixiang Ye
摘要
Discrete Diffusion Models (DDMs) have shown great potential in image generation, especially when equipped with reinforcement learning (RL) techniques. However, a fundamental yet overlooked limitation is revealed in our experiments: severe imbalance of credit assignment across timesteps during training. As a result, early generation timesteps, which carry higher exploration potential and determine the global structure, provide a smaller contribution to policy optimization. To conquer this, we propose a simple-yet-effective approach, to Re-balance timestep credit of Reinforcement Learning (RebRL) for better explorationexploitation trade-off and more efficient training of DDMs. RebRL is plug-and-play-simply replacing uniform temporal policy with strategic rebalancing along masking stages. RebRL is analytically plausible-derivation and analysis show that it enjoys a uniform token-level policy gradient, which benefits policy optimization. Experiments on textto-image generation benchmarks show that RebRL achieves state-of-the-art performance on GenEval and improves human preference score by up to 3.40 while effectively reducing training steps by ∼40%.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
相关 Paper
- Distillation Models are Good Samplers for Diffusion Reinforcement LearningZunxu Liu, Aiqiu Wu, Zhaofan Qiu, Yingwei Pan 等ICML 2026
- Stepwise Credit Assignment for GRPO on Flow-Matching ModelsYash Savani, Branislav Kveton, Yuchen Liu, Yilin Wang 等CVPR 2026 · 被引用 11 次
- PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation ModelsJeongjae Lee, Jong Chul YeICLR 2026 · 被引用 2 次
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationYifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun 等NeurIPS 2025 · 被引用 12 次
- UDM-GRPO: Stable and Efficient Group Relative Policy Optimization for Uniform Discrete Diffusion ModelsJiaqi Wang, Haoge Deng, Ting Pan, Yang Liu 等ICML 2026 · 被引用 3 次
