DiverseGRPO: Mitigating Mode Collapse in Image Generation via Diversity-Aware GRPO
Henglin Liu, Huijuan Huang, Jing Wang, Chang Liu, Xiu Li, Xiangyang Ji
Abstract
Reinforcement learning (RL), particularly GRPO, improves image generation quality significantly by comparing the relative performance of images generated within the same group. However, in the later stages of training, the model tends to produce homogenized outputs, lacking creativity and visual diversity, restricting the application scenarios of the model.This issue can be analyzed from both reward modeling and generation dynamics perspectives. First, traditional GRPO relies on single-sample quality as the reward signal, driving the model to converge toward a few high-reward generation modes while neglecting distribution-level diversity. Second, conventional GRPO regularization neglects the dominant role of early-stage denoising in preserving diversity, causing a misaligned regularization budget that limits the achievable quality–diversity trade-off.Motivated by these insights, we revisit the diversity degradation problem from both reward modeling and generation dynamics. At the reward level, we propose a distributional creativity bonus based on semantic grouping. Specifically, we construct a distribution-level representation via spectral clustering over samples generated from the same caption, and adaptively allocate exploratory rewards according to group sizes to encourage the discovery of novel visual modes. At the generation level, we introduce a structure-aware regularization, which enforces stronger early-stage constraints to preserve diversity without compromising reward optimization efficiency. Experiments demonstrate that our method achieves an 13%18% improvement in semantic diversity under matched quality scores, establishing a new Pareto frontier between image quality and diversity for GRPO-based image generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 62c71338-4f9a-4bdb-9d8a-84c0ef893daaCited by top-tier papers3
- Thinking with Frames: Generative Video Distortion Evaluation via Frame Reward ModelYuan Wang, Borui Liao, Huijuan Huang, Jinda Lu et al.CVPR 2026 · 5 citations
- DRM: Diffusion-based Reward Model With Step-wise GuidanceJaxon Zhang, Binxin Yang, Hubery Yin, Chen Li et al.CVPR 2026 · 1 citation
- Embedding-perturbed Exploration Preference Optimization for Flow ModelsSujie Hu, Chubin Chen, Jiashu Zhu, Jiahong Wu et al.ICML 2026
Builds on13
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- CLIPScore: A Reference-free Evaluation Metric for Image CaptioningJack Hessel, Ari Holtzman, Maxwell Forbes, Ronan Le Bras et al.EMNLP 2021 · 937 citations
Related papers
- From Competition to Synergy: Unlocking Reinforcement Learning for Subject-Driven Image GenerationZiwei Huang, Yin Shu, Hao Fang, Quanyu Long et al.ACL 2026 · 4 citations
- Seeing What Matters: Visual Preference Policy Optimization for Visual GenerationZiqi Ni, Yuanzhi Liang, Rui Li, Yi Zhou et al.CVPR 2026 · 9 citations
- Adaptive Divergence Regularized Policy Optimization for Fine-tuning Generative ModelsJiajun Fan, Tong Wei, Chaoran Cheng, Yuxin Chen et al.NeurIPS 2025 · 7 citations
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPOChengzhuo Tong, Ziyu Guo, Renrui Zhang, Wenyu Shan et al.NeurIPS 2025 · 39 citations
- From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image GenerationHan Song, Yucheng Zhou, Jianbing Shen, Yu ChengICLR 2026 · 9 citations
