Stepwise Credit Assignment for GRPO on Flow-Matching Models
Yash Savani, Branislav Kveton, Yuchen Liu, Yilin Wang, Jing Shi, Subhojyoti Mukherjee, Nikos Vlassis, Krishna Kumar Singh
Abstract
Flow-GRPO successfully applies reinforcement learning to flow models, but uses uniform credit assignment across all steps. This ignores the temporal structure of diffusion generation: early steps determine composition and content (low-frequency structure), while late steps resolve details and textures (high-frequency details). Moreover, assigning uniform credit based solely on the final image can inadvertently reward suboptimal intermediate steps, especially when errors are corrected later in the diffusion trajectory. We propose Stepwise-Flow-GRPO, which assigns credit based on each step's reward improvement. By leveraging Tweedie's formula to obtain intermediate reward estimates and introducing gain-based advantages, our method achieves superior sample efficiency and faster convergence. We also introduce a DDIM-inspired SDE that improves reward quality while preserving stochasticity for policy gradients.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c66c4138-518e-40ac-adf3-4fdc0f19bb9dBuilds on22
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- iGRPO: Fast Online RL for Flow Matching Model with Instant RewardSucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song et al.ICML 2026
- Fine-Grained GRPO for Precise Preference Alignment in Flow ModelsYujie Zhou, Pengyang Ling, Jiazi Bu, Yibin Wang et al.CVPR 2026 · 19 citations
- TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELSXiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li et al.ICLR 2026 · 98 citations
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion ModelsZheng Ding, Weirui YeICLR 2026 · 29 citations
- DRM: Diffusion-based Reward Model With Step-wise GuidanceJaxon Zhang, Binxin Yang, Hubery Yin, Chen Li et al.CVPR 2026 · 1 citation
