Lune

CVPR2026顶会

Curriculum Group Policy Optimization: Adaptive Sampling for Unleashing the Potential of Text-to-Image Generation

Baoteng Li, Xianghao Zang, Xinran Wang, Xiangyu Na, Zhixiang He, Hao Sun, Chi Zhang, Zhongjiang He, Tianwei Cao, Kongming Liang, Zhanyu Ma

2026年份

摘要

Text-to-Image (T2I) generation technology has achieved remarkable progress in recent years. Concurrently, reinforcement learning methods, particularly those based on Group Relative Policy Optimization (GRPO), have attracted widespread attention and have been successfully applied to T2I tasks. However, the uniform sampling strategy commonly adopted during training often ignores the match between sample difficulty and the model’s current learning capability, leading to low training efficiency. We argue that the key to unleashing the model’s potential lies in continuously providing ``high-value samples'' that match its evolving competence. To this end, we propose Curriculum Group Policy Optimization (CGPO), an adaptive curriculum training framework. During training, each prompt is used to generate a group of images, and a reward model assigns a reward to each image. We use the variance of these rewards as a proxy indicator—higher variance implies the model's understanding of the prompt is still unstable, indicating stronger learnability and thus higher value. CGPO adaptively constructs the curriculum by dynamically identifying and selecting high-value samples for training based on reward variance. Additionally, to address data imbalance in multi-category datasets, we design a category calibration method based on proportional fairness optimization, which balances training difficulty across categories. Experiments on GenEval, T2I-CompBench++, and DPG Bench demonstrate that our framework effectively improves generation performance.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper30

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖