GDRO: Group-level Reward Post-training Suitable for Diffusion Models
Yiyang Wang, Xi Chen, Xiaogang Xu, Yu Liu, Hengshuang Zhao
摘要
Recent advancements adopt online reinforcement learning (RL) from LLMs to text-to-image rectified flow diffusion models for reward alignment. The use of group-level rewards successfully aligns the model with the targeted reward. However, it faces challenges including low efficiency, dependency on stochastic samplers, and reward hacking. The problem is that rectified flow models are fundamentally different from LLMs: 1) For efficiency, online image sampling takes much more time and dominates the time of training. 2) For stochasticity, rectified flow is deterministic once the initial noise is fixed. Aiming at these problems and inspired by the effects of group-level rewards from LLMs, we design Group-level Direct Reward Optimization (GDRO). GDRO is a new post-training paradigm for group-level reward alignment that combines the characteristics of rectified flow models. Through rigorous theoretical analysis, we point out that GDRO supports full offline training that saves the large time cost for image rollout sampling. Also, it is diffusion-sampler-independent, which eliminates the need for the ODE-to-SDE approximation to obtain stochasticity. We also empirically study the reward hacking trap that may mislead the evaluation, and involve this factor in the evaluation using a corrected score that not only considers the original evaluation reward but also the trend of reward hacking. Extensive experiments demonstrate that GDRO effectively and efficiently improves the reward score of the diffusion model through group-wise offline optimization across the OCR and GenEval tasks, while demonstrating strong stability and robustness in mitigating reward hacking.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Reinforcing Diffusion Models by Direct Group Preference OptimizationYihong Luo, Tianyang Hu, Jing TangICLR 2026 · 被引用 13 次
- Rethinking DPO-Style Diffusion Aligning FrameworksXun Wu, Shaohan Huang, Lingjie Jiang, Furu WeiICCV 2025 · 被引用 4 次
- Inference-Time Alignment of Diffusion Models with Direct Noise OptimizationZhiwei Tang, Jiangweizhi Peng, Jiasheng Tang, Mingyi Hong 等ICML 2025
- iGRPO: Fast Online RL for Flow Matching Model with Instant RewardSucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song 等ICML 2026
- Reward Sharpness-Aware Fine-Tuning for Diffusion ModelsKwanyoung Kim, Byeongsu SimCVPR 2026 · 被引用 1 次
