PeRL: Permutation-Enhanced Reinforcement Learning for Interleaved Vision-Language Reasoning
Yizhen Zhang, Yang Ding, Shuoshuo Zhang, Xinchen Zhang, Haoling Li, Zhong-Zhi Li, Peijie Wang, Jie Wu, Lei Ji, Yeyun Gong, Yelong Shen, Yujiu Yang
摘要
Inspired by the impressive reasoning capabilities demonstrated by reinforcement learning approaches like DeepSeek-R1, recent emerging research has begun exploring the use of reinforcement learning (RL) to enhance vision-language models (VLMs) for multimodal reasoning tasks. However, most existing multimodal reinforcement learning approaches remain limited to spatial reasoning within single-image contexts, yet still struggle to generalize to more complex and real-world scenarios involving multi-image positional reasoning, where understanding the relationships across images is crucial. To address this challenge, we propose a general reinforcement learning approach PeRL tailored for interleaved multimodal tasks, and a multi-stage strategy designed to enhance the exploration-exploitation trade-off, thereby improving learning efficiency and task performance. Specifically, we introduce permutation of image sequences to simulate varied positional relationships to explore more spatial and positional diversity. Furthermore, we design a rollout filtering mechanism for resampling to focus on trajectories that contribute most to learning optimal behaviors to exploit learned policies effectively. We evaluate our model on 5 widely-used multi-image benchmarks and 3 single-image benchmarks. Our experiments confirm that PeRL trained model consistently surpasses R1-related and interleaved VLM baselines by a large margin, achieving state-of-the-art performance on multi-image benchmarks, while preserving comparable performance on single-image tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- UniCTokens: Boosting Personalized Understanding and Generation via Unified Concept TokensRuichuan An, Sihan Yang, Renrui Zhang, Zijun Shen 等NeurIPS 2025 · 被引用 61 次
- Unlocking Multimodal Mathematical Reasoning via Process Reward ModelRuilin Luo, Zhuofan Zheng, Lei Wang, Yifan Wang 等NeurIPS 2025 · 被引用 38 次
- HSSBench: Benchmarking Humanities and Social Sciences Ability for Multimodal Large Language ModelsZhaolu Kang, Junhao Gong, Jiaxu Yan, Wanke Xia 等ICLR 2026 · 被引用 24 次
- Generative Universal Verifier as Multimodal Meta-ReasonerXinchen Zhang, Xiaoying Zhang, Youbin Wu, Yanbin Cao 等ICLR 2026 · 被引用 20 次
- HermesFlow: Seamlessly Closing the Gap in Multimodal Understanding and GenerationLing Yang, Xinchen Zhang, Ye Tian, Shiyi Zhang 等NeurIPS 2025 · 被引用 16 次
它引用的顶会 Paper30
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech 等NeurIPS 2022 · 被引用 6,707 次
- MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual ContextsPan Lu, Hritik Bansal, Tony Xia, Jiacheng Liu 等ICLR 2024 · 被引用 1,472 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
- Visual-RFT: Visual Reinforcement Fine-TuningZiyu Liu, Zeyi Sun, Yuhang Zang, Xiaoyi Dong 等ICCV 2025 · 被引用 563 次
- VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement LearningHaozhe Wang, Chao Qu, Zuming Huang, Wei Chu 等NeurIPS 2025 · 被引用 356 次
相关 Paper
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesYihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng 等NeurIPS 2025 · 被引用 61 次
- Retrv-R1: A Reasoning-Driven MLLM Framework for Universal and Efficient Multimodal RetrievalLanyun Zhu, Deyi Ji, Tianrun Chen, Haiyang Wu 等NeurIPS 2025 · 被引用 12 次
- Incentivizing Versatile Video Reasoning in MLLMs via Data-Efficient Reinforcement LearningXiaodong Wang, Zhirong Wu, Langling Huang, Yuxi Zheng 等CVPR 2026
- VisPlay: Self-Evolving Vision-Language ModelsYicheng He, Chengsong Huang, Zongxia Li, Jiaxin Huang 等CVPR 2026 · 被引用 3 次
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception RewardTong Xiao, Xin Xu, Zhenya Huang, Hongyu Gao 等ICLR 2026 · 被引用 33 次
