Fine-Grained GRPO for Precise Preference Alignment in Flow Models
Yujie Zhou, Pengyang Ling, Jiazi Bu, Yibin Wang, Yuhang Zang, Jiaqi Wang, Li Niu, Guangtao Zhai
摘要
The incorporation of online reinforcement learning (RL) into diffusion and flow-based generative models has recently gained attention as a powerful paradigm for aligning model behavior with human preferences. By leveraging stochastic sampling via Stochastic Differential Equations (SDEs) during the denoising phase, these models can explore a variety of denoising trajectories, enhancing the exploratory capacity of RL. However, despite their ability to discover potentially high-reward samples, current approaches often struggle to effectively align with preferences due to the sparsity and narrowness of reward feedback. To overcome this limitation, we introduce a novel framework called Granular-GRPO (GRPO), which enables fine-grained and comprehensive evaluation of sampling directions in the RL training of flow models. Specifically, we propose a Singular Stochastic Sampling mechanism that supports step-wise stochastic exploration while ensuring strong correlation between injected noise and reward signals, enabling more accurate credit assignment to each SDE perturbation. Additionally, to mitigate the bias introduced by fixed-granularity denoising, we design a Multi-Granularity Advantage Integration module that aggregates advantages computed across multiple diffusion scales, resulting in a more robust and holistic assessment of sampling trajectories. Extensive experiments on various reward models, including both in-domain and out-of-domain settings, demonstrate that our GRPO outperforms existing flow-based GRPO baselines, highlighting its effectiveness and generalization capability.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- GRPO-Guard: Mitigating Implicit Over-Optimization in Flow Matching via Regulated ClippingJing Wang, Jiajun Liang, Jie Liu, Henglin Liu 等CVPR 2026 · 被引用 46 次
- Stepwise Credit Assignment for GRPO on Flow-Matching ModelsYash Savani, Branislav Kveton, Yuchen Liu, Yilin Wang 等CVPR 2026 · 被引用 11 次
- LeapAlign: Post-training Flow Matching Models at Any Generation Step by Building Two-Step TrajectoriesZhanhao Liang, Tao Yang, Jie Wu, Chengjian Feng 等CVPR 2026 · 被引用 6 次
- Talk2Move: Reinforcement Learning for Text-Instructed Object-Level Geometric Transformation in ScenesJing Tan, Zhaoyang Zhang, Yantao Shen, Jiarui Cai 等CVPR 2026 · 被引用 3 次
- Enhancing Spatial Understanding in Image Generation via Reward ModelingZhenyu Tang, Chaoran Feng, Yufan Deng, Jie Wu 等CVPR 2026 · 被引用 2 次
它引用的顶会 Paper27
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 被引用 11,724 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
相关 Paper
- DenseGRPO: From Sparse to Dense Reward for Flow Matching Model AlignmentHaoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang 等ICLR 2026 · 被引用 21 次
- iGRPO: Fast Online RL for Flow Matching Model with Instant RewardSucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song 等ICML 2026
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li 等NeurIPS 2025 · 被引用 647 次
- TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELSXiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li 等ICLR 2026 · 被引用 98 次
- TreeGRPO: Tree-Advantage GRPO for Online RL Post-Training of Diffusion ModelsZheng Ding, Weirui YeICLR 2026 · 被引用 29 次
