Janus-Pro-R1: Advancing Collaborative Visual Comprehension and Generation via Reinforcement Learning
Kaihang Pan, Yang Wu, Wendong Bu, Kai Shen, Juncheng Li, Yingting Wang, Yunfei Li, Siliang Tang, Jun Xiao, Fei Wu, Hang Zhao, Yueting Zhuang
Abstract
Recent endeavors in Multimodal Large Language Models (MLLMs) aim to unify visual comprehension and generation. However, these two capabilities remain largely independent, as if they are two separate functions encapsulated within the same model. Consequently, visual comprehension does not enhance visual generation, and the reasoning mechanisms of LLMs have not been fully integrated to revolutionize image generation. In this paper, we propose to enable the collaborative co-evolution of visual comprehension and generation, advancing image generation into an iterative introspective process. We introduce a two-stage training approach: supervised fine-tuning teaches the MLLM with the foundational ability to generate genuine CoT for visual generation, while reinforcement learning activates its full potential via an exploration-exploitation trade-off. Ultimately, we unlock the Aha moment in visual generation, advancing MLLMs from text-to-image tasks to unified image generation. Extensive experiments demonstrate that our model not only excels in text-to-image generation and image editing, but also functions as a superior image semantic evaluator with enhanced visual comprehension capabilities. Project Page: https://janus-pro-r1.github.io.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 34a509b5-15c2-4b1c-a90b-de326f9ef1d3Cited by top-tier papers5
- WiseEdit: Benchmarking Cognition- and Creativity-Informed Image EditingKaihang Pan, Weile Chen, Haiyi Qiu, Qifan Yu et al.CVPR 2026 · 9 citations
- Residual Decoder Adapter: ID-Preserving Tokenizer Adaption for Autoregressive Text RenderingDongxing Mao, Alex Jinpeng Wang, Jiahao Tang, Kevin Qinghong Lin et al.CVPR 2026 · 1 citation
- Benchmarking and Improving Fine-Grained Text-to-Image Alignment via Paired Reinforcement LearningKaihang Pan, Wendong Bu, Yuruo Wu, Kai Shen et al.ICML 2026
- MapDream: Task-Driven Map Learning for Vision-Language NavigationGuoxin Lian, Shuo Wang, Yucheng Wang, Yongcai Wang et al.ICML 2026
- OSPO: Object-Centric Self-Improving Preference Optimization for Text-to-Image GenerationYoonjin Oh, Yongjin Kim, Hyomin Kim, Donghwan Chi et al.CVPR 2026
Builds on33
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang et al.NeurIPS 2025 · 15 citations
- ThinkGen: Generalized Thinking for Visual GenerationSiyu Jiao, Yiheng Lin, Yujie Zhong, Qi She et al.CVPR 2026 · 12 citations
- Generative Multimodal Pretraining with Discrete Diffusion Timestep TokensKaihang Pan, Wang Lin, Zhongqi Yue, Tenglong Ao et al.CVPR 2025
- Auto-Encoding Morph-Tokens for Multimodal LLMKaihang Pan, Siliang Tang, Juncheng Li, Zhaoyu Fan et al.ICML 2024 · 36 citations
- T2I-R1: Reinforcing Image Generation with Collaborative Semantic-level and Token-level CoTDongzhi Jiang, Ziyu Guo, Renrui Zhang, Zhuofan Zong et al.NeurIPS 2025 · 181 citations
