Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception Reward
Tong Xiao, Xin Xu, Zhenya Huang, Hongyu Gao, Quan Liu, Qi Liu, Enhong Chen
摘要
Enhancing the multimodal reasoning capabilities of Multimodal Large Language Models (MLLMs) is a challenging task that has attracted increasing attention in the community. Recently, several studies have applied Reinforcement Learning with Verifiable Rewards (RLVR) to the multimodal domain in order to enhance the reasoning abilities of MLLMs. However, these works largely overlook the enhancement of multimodal perception capabilities in MLLMs, which serve as a core prerequisite and foundational component of complex multimodal reasoning. Through McNemar's test, we find that existing RLVR method fails to effectively enhance the multimodal perception capabilities of MLLMs, thereby limiting their further improvement in multimodal reasoning. To address this limitation, we propose Perception-R1, which introduces a novel visual perception reward that explicitly encourages MLLMs to perceive the visual content accurately, thereby can effectively incentivizing both their multimodal perception and reasoning capabilities. Specifically, we first collect textual visual annotations from the CoT trajectories of multimodal problems, which will serve as visual references for reward assignment. During RLVR training, we employ a judging LLM to assess the consistency between the visual annotations and the responses generated by MLLM, and assign the visual perception reward based on these consistency judgments. Extensive experiments on several multimodal math and general benchmarks demonstrate the effectiveness and robustness of our Perception-R1, which achieves superior performance on all benchmarks using only 1,442 training data. Our code and dataset will be available at https://github.com/tongxiao2002/Perception-R1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal ReasoningChi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu 等CVPR 2026 · 被引用 22 次
- Thinking in Uncertainty: Mitigating Hallucinations in MLRMs with Latent Entropy-Aware DecodingZhongxing Xu, Zhonghua Wang, Zhe Qian, Dachuan Shi 等CVPR 2026 · 被引用 16 次
- APPO: Attention-guided Perception Policy Optimization for Video ReasoningHenghui Du, Chang Zhou, Xi Chen, Di HuCVPR 2026 · 被引用 3 次
- PDCR: Perception-Decomposed Confidence Reward for Vision-Language ReasoningHee Suk Yoon, Eunseop Yoon, Ji Woo Hong, SooHwan Eom 等CVPR 2026 · 被引用 3 次
- Children's Intelligence Tests Pose Challenges for MLLMs? KidGym: A 2D Grid-Based Reasoning Benchmark for MLLMsHengwei Ye, Yuanting Guan, Yuxuan Ge, Tianying Zhu 等ICLR 2026 · 被引用 2 次
它引用的顶会 Paper24
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- InstructBLIP: Towards General-purpose Vision-Language Models with Instruction TuningWenliang Dai, Junnan Li, Dongxu Li, Anthony Meng Huat Tiong 等NeurIPS 2023 · 被引用 4,013 次
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan 等NeurIPS 2025 · 被引用 2,828 次
- Are We on the Right Way for Evaluating Large Vision-Language Models?Lin Chen, Jinsong Li, Xiaoyi Dong, Pan Zhang 等NeurIPS 2024 · 被引用 1,029 次
相关 Paper
- Spotlight on Token Perception for Multimodal Reinforcement LearningSiyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo 等ICLR 2026 · 被引用 45 次
- Cantor: Inspiring Multimodal Chain-of-Thought of MLLMTimin Gao, Peixian Chen, Mengdan Zhang, Chaoyou Fu 等ACM MM 2024 · 被引用 20 次
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningZhangquan Chen, Xufang Luo, Dongsheng LiICCV 2025 · 被引用 2 次
- VisuRiddles: Fine-grained Perception is a Primary Bottleneck for Multimodal Large Language Models in Abstract Visual ReasoningHao Yan, Xingchen Liu, Hao Wang, Zhenbiao Cao 等ICLR 2026 · 被引用 7 次
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye 等ICLR 2026 · 被引用 670 次
