GoT-R1: Unleashing Reasoning Capability of Autoregressive Visual Generation with Reinforcement Learning
Chengqi Duan, Rongyao Fang, Yuqing Wang, Kun Wang, Linjiang Huang, Xingyu Zeng, Hongsheng Li, Xihui Liu
Abstract
Visual generation models have made remarkable progress in creating realistic images from text prompts, yet struggle with complex prompts that specify multiple objects with precise spatial relationships and attributes. Effective handling of such prompts requires explicit reasoning about the semantic content and spatial layout. We present GoT-R1, a framework that applies reinforcement learning to enhance semantic-spatial reasoning in autoregressive visual generation models. Leveraging the natural affinity between autoregressive architectures and sequential reasoning, our approach builds upon the Generation Chain-of-Thought framework to enable models to autonomously discover effective reasoning strategies beyond predefined templates. To achieve this, we propose a dual-stage multi-dimensional reward framework that leverages MLLMs to evaluate both the reasoning process and final output, enabling effective supervision across the entire generation pipeline. The reward system assesses semantic alignment, spatial accuracy, and visual quality in a unified approach. Experimental results demonstrate significant improvements on T2I-CompBench and GenEval benchmark, particularly in compositional tasks involving precise spatial relationships and attribute binding. GoT-R1 advances the state-of-the-art in autoregressive image generation by successfully transferring sophisticated reasoning capabilities from language models to the visual generation domain. Code is available at https://github.com/gogoduan/GoT-R1.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d9ed942e-323c-4fd0-bb2b-ef754dc9c6aaCited by top-tier papers7
- JarvisEvo: Towards a Self-Evolving Photo Editing Agent with Synergistic Editor-Evaluator OptimizationYunlong Lin, Linqing Wang, Kunjie Lin, Zixu Lin et al.CVPR 2026 · 31 citations
- PSR: Scaling Multi-Subject Personalized Image Generation with Pairwise Subject-Consistency RewardsShulei Wang, Longhui Wei, Xin He, Jianbo Ouyang et al.CVPR 2026 · 7 citations
- Learning to Generate via Understanding: Understanding-Driven Intrinsic Rewarding for Unified Multimodal ModelsJiadong Pan, Liang Li, Yuxin Peng, Yu-Ming Tang et al.CVPR 2026 · 5 citations
- Meta-CoT: Enhancing Granularity and Generalization in Image EditingShiyi Zhang, Yiji Cheng, Tiankai Hang, Zijin Yin et al.CVPR 2026 · 3 citations
- Saliency-R1: Enforcing Interpretable and Faithful Vision-language Reasoning via Saliency-map Alignment RewardShizhan Gong, Minda Hu, Qiyuan Zhang, Chen Ma et al.CVPR 2026 · 1 citation
Builds on22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
Related papers
- RePrompt: Reasoning-Augmented Reprompting for Text-to-Image Generation via Reinforcement LearningMingrui Wu, Lu Wang, Pu Zhao, Fangkai Yang et al.ICLR 2026 · 19 citations
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
- GoT: Unleashing Reasoning Capability of MLLM for Visual Generation and EditingRongyao Fang, Chengqi Duan, Kun Wang, Linjiang Huang et al.NeurIPS 2025 · 5 citations
- SpatialReward: Verifiable Spatial Reward Modeling for Fine-Grained Spatial Consistency in Text-to-Image GenerationSashuai zhou, Qiang Zhou, Ma Junpeng, Yue Cao et al.CVPR 2026 · 7 citations
- ThinkGen: Generalized Thinking for Visual GenerationSiyu Jiao, Yiheng Lin, Yujie Zhong, Qi She et al.CVPR 2026 · 12 citations
