Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
Ming Nie, Chunwei Wang, Jianhua Han, Hang Xu, Li Zhang
Abstract
Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks like visual storytelling and step-by-step visual reasoning. In this work, we propose a reinforcement learning-based post-training strategy to unlock this capability in existing unified models, without relying on large-scale multimodal interleaved datasets. We begin with a warm-up stage using a hybrid dataset comprising curated interleaved sequences and limited data for multimodal understanding and text-to-image generation, which exposes the model to interleaved generation patterns while preserving its pretrained capabilities. To further refine interleaved generation, we propose a unified policy optimization framework that extends Group Relative Policy Optimization (GRPO) to the multimodal setting. Our approach jointly models text and image generation within a single decoding trajectory and optimizes it with our novel hybrid rewards covering textual relevance, visual-text alignment, and structural fidelity. Additionally, we incorporate process-level rewards to provide step-wise guidance, enhancing training efficiency in complex multimodal tasks. Experiments on MMIE and InterleavedBench demonstrate that our approach significantly enhances the quality and coherence of multimodal interleaved generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 85274a04-5feb-45c5-916c-4014cc602ddeCited by top-tier papers2
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in Unified Multimodal Models via Decompositional Verifiable RewardRunhui Huang, Jie Wu, Rui Yang, Zhe Liu et al.ICML 2026
- How RL Unlocks the Aha Moment in Geometric Interleaved ReasoningXiangxiang Zhang, Caijun jia, Siyuan Li, he dingyu et al.ICML 2026
Builds on19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong et al.NeurIPS 2023 · 1,310 citations
Related papers
- Unified Multimodal Models as Auto-EncodersZhiyuan Yan, Kaiqing Lin, Zongjian Li, Junyan Ye et al.CVPR 2026 · 12 citations
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang et al.NeurIPS 2025 · 15 citations
- EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction EvolutionZhebei Shen, Qifan Yu, Juncheng Li, Wei Ji et al.NeurIPS 2025 · 2 citations
- VisPlay: Self-Evolving Vision-Language ModelsYicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang et al.CVPR 2026 · 3 citations
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image GenerationQian Liang, Yujia Wu, Kuncheng Li, Jiwei Wei et al.AAAI 2026 · 6 citations
