MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image Generation
Qian Liang, Yujia Wu, Kuncheng Li, Jiwei Wei, Shiyuan He, Jinyu Guo, Ning Xie
Abstract
Multimodal Large Language Models (MLLMs) with unified architectures excel across a wide range of vision-language tasks, yet aligning them with personalized image generation remains a significant challenge. Existing methods for MLLMs are frequently subject-specific, demanding a data-intensive fine-tuning process for every new subject, which limits their scalability. In this paper, we introduce MM-R1, a framework that integrates a cross-modal Chain-of-Thought (X-CoT) reasoning strategy to unlock the inherent potential of unified MLLMs for personalized image generation. Specifically, we structure personalization as an integrated visual reasoning and generation process: (1) grounding subject concepts by interpreting and understanding user-provided images and contextual cues, and (2) generating personalized images conditioned on both the extracted subject representations and user prompts. To further enhance the reasoning capability, we adopt Grouped Reward Proximal Policy Optimization(GRPO) to explicitly align the generation. Experiments demonstrate that MM-R1 unleashes the personalization capability of unified MLLMs to generate images with high subject fidelity and strong text alignment in a zero-shot manner.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi et al.CVPR 2026 · 16 citations
- DCoAR: Deep Concept Injection into Unified Autoregressive Models for Personalized Text-to-Image GenerationFangtai Wu, Mushui Liu, Weijie He, Zhao Wang et al.CVPR 2026 · 1 citation
Builds on21
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Large Language Models are Zero-Shot ReasonersTakeshi Kojima, Shixiang Shane Gu, Machel Reid, Yutaka Matsuo et al.NeurIPS 2022 · 8,168 citations
Related papers
- UniMo: Unified Motion Generation and Understanding with Chain of ThoughtGuocun Wang, Kenkun Liu, Jing Lin, Guorui Song et al.AAAI 2026
- Towards Unified Multimodal Interleaved Generation via Group Relative Policy OptimizationMing Nie, Chunwei Wang, Jianhua Han, Hang Xu et al.NeurIPS 2025 · 7 citations
- ThinkGen: Generalized Thinking for Visual GenerationSiyu Jiao, Yiheng Lin, Yujie Zhong, Qi She et al.CVPR 2026 · 12 citations
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang et al.NeurIPS 2025 · 15 citations
- Unified Multimodal Chain-of-Thought Reward Model through Reinforcement Fine-TuningYibin Wang, Zhimin Li, Yuhang Zang, Chunyu Wang et al.NeurIPS 2025 · 102 citations
