Towards Unified Multimodal Interleaved Generation via Group Relative Policy Optimization
Ming Nie, Chunwei Wang, Jianhua Han, Hang Xu, Li Zhang
摘要
Unified vision-language models have made significant progress in multimodal understanding and generation, yet they largely fall short in producing multimodal interleaved outputs, which is a crucial capability for tasks like visual storytelling and step-by-step visual reasoning. In this work, we propose a reinforcement learning-based post-training strategy to unlock this capability in existing unified models, without relying on large-scale multimodal interleaved datasets. We begin with a warm-up stage using a hybrid dataset comprising curated interleaved sequences and limited data for multimodal understanding and text-to-image generation, which exposes the model to interleaved generation patterns while preserving its pretrained capabilities. To further refine interleaved generation, we propose a unified policy optimization framework that extends Group Relative Policy Optimization (GRPO) to the multimodal setting. Our approach jointly models text and image generation within a single decoding trajectory and optimizes it with our novel hybrid rewards covering textual relevance, visual-text alignment, and structural fidelity. Additionally, we incorporate process-level rewards to provide step-wise guidance, enhancing training efficiency in complex multimodal tasks. Experiments on MMIE and InterleavedBench demonstrate that our approach significantly enhances the quality and coherence of multimodal interleaved generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- AlphaGRPO: Unlocking Self-Reflective Multimodal Generation in Unified Multimodal Models via Decompositional Verifiable RewardRunhui Huang, Jie Wu, Rui Yang, Zhe Liu 等ICML 2026
- How RL Unlocks the Aha Moment in Geometric Interleaved ReasoningXiangxiang Zhang, Caijun jia, Siyuan Li, he dingyu 等ICML 2026
它引用的顶会 Paper19
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann 等ICLR 2024 · 被引用 4,569 次
- ImageReward: Learning and Evaluating Human Preferences for Text-to-Image GenerationJiazheng Xu, Xiao Liu, Yuchen Wu, Yuxuan Tong 等NeurIPS 2023 · 被引用 1,310 次
相关 Paper
- Unified Multimodal Models as Auto-EncodersZhiyuan Yan, Kaiqing Lin, Zongjian Li, Junyan Ye 等CVPR 2026 · 被引用 12 次
- Co-Reinforcement Learning for Unified Multimodal Understanding and GenerationJingjing Jiang, Chongjie Si, Jun Luo, Hanwang Zhang 等NeurIPS 2025 · 被引用 15 次
- EvolvedGRPO: Unlocking Reasoning in LVLMs via Progressive Instruction EvolutionZhebei Shen, Qifan Yu, Juncheng Li, Wei Ji 等NeurIPS 2025 · 被引用 2 次
- VisPlay: Self-Evolving Vision-Language ModelsYicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang 等CVPR 2026 · 被引用 3 次
- MM-R1: Unleashing the Power of Unified Multimodal Large Language Models for Personalized Image GenerationQian Liang, Yujia Wu, Kuncheng Li, Jiwei Wei 等AAAI 2026 · 被引用 6 次
