Alleviating Sparse Rewards by Modeling Step-Wise and Long-Term Sampling Effects in Flow-Based GRPO
Yunze Tong, Mushui Liu, Canyu Zhao, Wanggui He, Shiyi Zhang, Peng Zhang, Hongwei Zhang, Jinlong Liu, Hao Jiang
摘要
Deploying GRPO on Flow Matching models has proven effective for text-to-image generation. However, existing paradigms typically propagate an outcome-based reward to all preceding denoising steps without distinguishing the local effect of each step. Moreover, current group-wise ranking mainly compares trajectories at matched timesteps and ignores within-trajectory dependencies, where certain early denoising actions can affect later states via delayed, implicit interactions. We propose TurningPoint-GRPO (TP-GRPO), a GRPO framework that alleviates stepwise reward sparsity and explicitly models longterm effects within the denoising trajectory. TP-GRPO makes two key innovations: (i) it replaces outcome-based rewards with step-level incremental rewards, providing a dense, step-aware learning signal that better isolates each denoising action's "pure" effect, and (ii) it identifies turning points-steps that flip the local reward trend and make subsequent reward evolution consistent with the overall trajectory trend-and assigns these actions an aggregated long-term reward to capture their delayed impact. Turning points are detected solely via sign changes in incremental rewards, making TP-GRPO efficient and hyperparameterfree. Extensive experiments also demonstrate that TP-GRPO exploits reward signals more effectively and consistently improves generation. Code is available at https://github.com/ YunzeTong/TurningPoint-GRPO .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper33
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Elucidating the Design Space of Diffusion-Based Generative ModelsTero Karras, Miika Aittala, Timo Aila, Samuli LaineNeurIPS 2022 · 被引用 3,959 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
相关 Paper
- TEMPFLOW-GRPO: WHEN TIMING MATTERS FOR GRPO IN FLOW MODELSXiaoxuan He, Siming Fu, Yuke Zhao, Wanli Li 等ICLR 2026 · 被引用 98 次
- iGRPO: Fast Online RL for Flow Matching Model with Instant RewardSucheng Ren, Chen Chen, Zhenbang Wang, Liangchen Song 等ICML 2026
- DenseGRPO: From Sparse to Dense Reward for Flow Matching Model AlignmentHaoyou Deng, Keyu Yan, Chaojie Mao, Xiang Wang 等ICLR 2026 · 被引用 21 次
- Stepwise Credit Assignment for GRPO on Flow-Matching ModelsYash Savani, Branislav Kveton, Yuchen Liu, Yilin Wang 等CVPR 2026 · 被引用 11 次
- Reinforcement Learning Meets Masked Generative Models: Mask-GRPO for Text-to-Image GenerationYifu Luo, Xinhao Hu, Keyu Fan, Haoyuan Sun 等NeurIPS 2025 · 被引用 12 次
