Spotlight on Token Perception for Multimodal Reinforcement Learning
Siyuan Huang, Xiaoye Qu, Yafu Li, Yun Luo, Zefeng He, Daizong Liu, Yu Cheng
Abstract
While Reinforcement Learning with Verifiable Rewards (RLVR) has advanced the reasoning capabilities of Large Vision-Language Models (LVLMs), most existing methods in multimodal reasoning neglect the critical role of visual perception within the RLVR optimization process. In this paper, we undertake a pioneering exploration of multimodal RLVR through the novel perspective of token perception, which measures the visual dependency of each generated token. With a granular analysis of Chain-of-Thought (CoT) processes, we uncover two key insights: first, token perception in a rollout trajectory is sparsely distributed, where only a small fraction of tokens have high visual dependency for visually-grounded reasoning; second, different trajectories exhibit significant divergence in their overall visual dependency. Based on these observations, we propose Visually-Perceptive Policy Optimization (VPPO), a novel policy gradient algorithm that explicitly leverages token perception to refine the learning signal. Specifically, VPPO achieves this through a dual mechanism: it reweights a trajectory's advantage by its overall visual dependency, and focuses policy updates exclusively on perceptually pivotal tokens. On a comprehensive suite of eight perception and reasoning benchmarks, VPPO demonstrates substantial gains over leading open-source RL-tuned models, with its effectiveness consistently validated across 7B and 32B model scales. Our findings not only establish a new token-level perceptual perspective for analyzing multimodal RLVR but also present a novel and effective optimization strategy to significantly enhance the multimodal reasoning capabilities of LVLMs. Our code is available at https://github.com/huaixuheqing/VPPO-RL .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7884c339-2bd1-40ef-9535-038c01d4ba84Cited by top-tier papers6
- DiffThinker: Towards Generative Multimodal Reasoning with Diffusion ModelsZefeng He, Xiaoye Qu, Yafu Li, Tong Zhu et al.ICML 2026 · 13 citations
- Visually-Guided Policy Optimization for Multimodal ReasoningZengbin Wang, Feng Xiong, Liang Lin, Xuecai Hu et al.ACL 2026 · 7 citations
- Decoupling Skeleton and Flesh: Efficient Multimodal Table Reasoning with Disentangled Alignment and Structure-aware GuidanceYingjie Zhu, Xuefeng Bai, Kehai Chen, Yang Xiang et al.ICML 2026 · 3 citations
- PDCR: Perception-Decomposed Confidence Reward for Vision-Language ReasoningHee Suk Yoon, Eunseop Yoon, Ji Woo Hong, SooHwan Eom et al.CVPR 2026 · 3 citations
- CARE What Fails: Contrastive Anchored-REflection for Verifiable Multimodal ReasoningYongxin Wang, Zhicheng Yang, Meng Cao, Mingfei Han et al.CVPR 2026
Builds on23
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- DAPO: An Open-Source LLM Reinforcement Learning System at ScaleQiying Yu, Zheng Zhang, Ruofei Zhu, Yufeng Yuan et al.NeurIPS 2025 · 2,828 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM ReasoningShenzhi Wang, Le Yu, Chang Gao, Chujie Zheng et al.NeurIPS 2025 · 592 citations
- MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding BenchmarkXiang Yue, Tianyu Zheng, Yuansheng Ni, Yubo Wang et al.ACL 2025 · 377 citations
Related papers
- Perception-Aware Policy Optimization for Multimodal ReasoningZhenhailong Wang, Xuehang Guo, Sofia Stoica, Haiyang Xu et al.ICLR 2026 · 104 citations
- More Thought, Less Accuracy? On the Dual Nature of Reasoning in Vision-Language ModelsXinyu Tian, Shu Zou, Zhaoyuan Yang, Mengqi He et al.ICLR 2026 · 29 citations
- Perception-R1: Advancing Multimodal Reasoning Capabilities of MLLMs via Visual Perception RewardTong Xiao, Xin Xu, Zhenya Huang, Hongyu Gao et al.ICLR 2026 · 33 citations
- CFPO: Counterfactual Policy Optimization for Multimodal ReasoningZhangyuan Yu, Wanran Sun, Guangjing Yang, Xiaohu Wu et al.ICML 2026
- Perceptual-Evidence Anchored Reinforced Learning for Multimodal ReasoningChi Zhang, Haibo Qiu, Qiming Zhang, Yufei Xu et al.CVPR 2026 · 22 citations
