PCPO: Proportionate Credit Policy Optimization for Preference Alignment of Image Generation Models
Jeongjae Lee, Jong Chul Ye
Abstract
While reinforcement learning has advanced the alignment of text-to-image (T2I) models, state-of-the-art policy gradient methods are still hampered by training instability and high variance, hindering convergence speed and compromising image quality. Our analysis identifies a key cause of this instability: disproportionate credit assignment, in which the mathematical structure of the generative sampler produces volatile and non-proportional feedback across timesteps. To address this, we introduce Proportionate Credit Policy Optimization (PCPO), a framework that enforces proportional credit assignment through a stable objective reformulation and a principled reweighting of timesteps. This correction stabilizes the training process, leading to significantly accelerated convergence and superior image quality. The improvement in quality is a direct result of mitigating model collapse, a common failure mode in recursive training. PCPO substantially outperforms existing policy gradient baselines on all fronts, including the state-of-the-art DanceGRPO. Code is available at https://github.com/jaylee2000/pcpo/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fe931513-9bac-4556-96d1-ff8d39af7f19Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
Related papers
- Stepwise Credit Assignment for GRPO on Flow-Matching ModelsYash Savani, Branislav Kveton, Yuchen Liu, Yilin Wang et al.CVPR 2026 · 11 citations
- VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive GenerationShikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu et al.CVPR 2026 · 2 citations
- From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image GenerationHan Song, Yucheng Zhou, Jianbing Shen, Yu ChengICLR 2026 · 9 citations
- Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image GenerationBaoteng Li, Xianghao Zang, Xinran Wang, Xiangyu Na et al.CVPR 2026
- Principled RL for Flow Matching Emerges from the Chunk-level Policy OptimizationYifu Luo, Haoyuan Sun, Xinhao Hu, Penghui Du et al.ICML 2026 · 12 citations
