Group Critical-token Policy Optimization for Autoregressive Image Generation
Guohui Zhang, Hu Yu, Xiaoxiao Ma, JingHao Zhang, Yaning Pan, Mingde Yao, Jie Xiao, Linjiang Huang, Jie Huang, Feng Zhao
Abstract
Recent studies have extended Reinforcement Learning with Verifiable Rewards (RLVR) to autoregressive (AR) visual generation and achieved promising progress. However, existing methods typically apply uniform optimization across all image tokens, while the varying contributions of different image tokens for RLVR's training remain unexplored. In fact, the key obstacle lies in how to identify more critical image tokens during AR generation and implement effective token-wise optimization for them. To tackle this challenge, we propose roup ritical-token olicy ptimization (), which facilitates effective policy optimization on critical tokens. We identify the critical tokens in RLVR-based AR generation from three perspectives, specifically: Causal dependency: early tokens fundamentally determine the later tokens and final image effect due to unidirectional dependency; Entropy-induced spatial structure: tokens with high entropy gradients correspond to image structure and bridges distinct visual regions; RLVR-focused token diversity: tokens with low visual similarity across a group of sampled images contribute to richer token-level diversity. For these identified critical tokens, we further introduce a dynamic token-wise advantage weight to encourage exploration, based on confidence divergence between the policy model and reference model. By leveraging 30% of the image tokens, GCPO achieves better performance than GRPO with full tokens. Extensive experiments on multiple text-to-image benchmarks for both AR models and unified multimodal models demonstrate the effectiveness of GCPO for AR visual generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- MaskFocus: Focusing Policy Optimization on Critical Steps for Masked Image GenerationGuohui Zhang, Hu Yu, Xiaoxiao Ma, Yaning Pan et al.CVPR 2026 · 6 citations
- VAR RL Done Right: Tackling Asynchronous Policy Conflicts in Visual Autoregressive GenerationShikun Sun, Liao Qu, Huichao Zhang, Yiheng Liu et al.CVPR 2026 · 2 citations
- ParaUni: Enhance Generation in Unified Multimodal Model with Reinforcement-driven Hierarchical Parallel Information InteractionJiangtong Tan, Lin Liu, Jie Huang, Xiaopeng Zhang et al.CVPR 2026 · 1 citation
- SJD-PAC: Accelerating Speculative Jacobi Decoding via Proactive Drafting and Adaptive ContinuationJialiang Kang, Han Shu, Wenshuo Li, Yingjie Zhai et al.CVPR 2026 · 1 citation
Builds on12
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
- Pick-a-Pic: An Open Dataset of User Preferences for Text-to-Image GenerationYuval Kirstain, Adam Polyak, Uriel Singer, Shahbuland Matiana et al.NeurIPS 2023 · 1,192 citations
- Flow-GRPO: Training Flow Matching Models via Online RLJie Liu, Gongye Liu, Jiajun Liang, Yangguang Li et al.NeurIPS 2025 · 647 citations
Related papers
- Visually-Guided Policy Optimization for Multimodal ReasoningZengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu et al.ACL 2026 · 7 citations
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo et al.ICLR 2026 · 45 citations
- From Broad Exploration to Stable Synthesis: Entropy-Guided Optimization for Autoregressive Image GenerationHan Song, Yucheng Zhou, Jianbing Shen, Yu ChengICLR 2026 · 9 citations
- Delving into RL for Image Generation with CoT: A Study on DPO vs. GRPOChengzhuo Tong, Ziyu Guo, Renrui Zhang, Wenyu Shan et al.NeurIPS 2025 · 39 citations
- GTPO and GRPO-S: Token and Sequence-Level Reward Shaping with Policy EntropyHongze Tan, Zihan Wang, Jianfei Pan, Jinghao Lin et al.ICML 2026 · 53 citations
