Seeing is Solving: Unlocking Efficient Multimodal RL via View Alignment
Qinsi Wang, Jing Shi, Kun Wan, Handong Zhao, Hancheng Ye, Zishan Shao, Jinghan Ke, Yudong Liu, Daniel Miranda, Purvak Lapsiya, Yiran Chen, Wentian Zhao
Abstract
Although Reinforcement Learning Fine-Tuning (RLFT) applied to Vision-Language Models (VLMs) substantially enhances multimodal reasoning capabilities, their prohibitive training cost limits broad adoption. Surprisingly, most existing methods simply port Large Language Model (LLM) RLFT techniques to VLMs, while ignoring a intrinsic property of multimodal models: their dynamic text–vision alignment. We ask a new question: Can this intrinsic alignment be turned into a training signal that makes VLM RLFT more efficient? We analyze how a VLM plans to attend, actually attends, and ideally should attend during reasoning, and derive two lightweight metrics from these patterns. Predictive View Accuracy (PVA) estimates sample difficulty, and Reasoning View Accuracy (RVA) reflects the quality of chain-of-thought (CoT) reasoning. These alignment signals enable automated data curriculum and dense reasoning supervision. We introduce FOCUS-RL, a plug-and-play framework that can be seamlessly integrated into any VLM and dramatically boosts RLFT training efficiency. FOCUS-RL achieves 2.5 x – 4 x faster convergence over vanilla GRPO and consistent accuracy gains (+4.4 on average) across six different benchmarks and multiple VLM families.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ee283d61-6258-457d-87bd-d22dbae964fcBuilds on28
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Optimizing Chain-of-Thought Reasoners via Gradient Variance Minimization in Rejection Sampling and RLJiarui Yao, Yifan Hao, Hanning Zhang, Hanze Dong et al.NeurIPS 2025 · 28 citations
- Reason-RFT: Reinforcement Fine-Tuning for Visual Reasoning of Vision Language ModelsHuajie Tan, Yuheng Ji, Xiaoshuai Hao, Xiansheng Chen et al.NeurIPS 2025 · 45 citations
- Point-RFT: Improving Multimodal Reasoning with Visually Grounded Reinforcement FinetuningMinheng Ni, Zhengyuan Yang, Linjie Li, Chung-Ching Lin et al.NeurIPS 2025 · 35 citations
- Chart-FR1: Visual Focus-Driven Fine-Grained Reasoning on Dense ChartsHongkun Pan, Yuwei Wu, Wanyi Hong, Shenghui Hu et al.CVPR 2026 · 1 citation
- VisRL: Intention-Driven Visual Perception via Reinforced ReasoningZhangquan Chen, Xufang Luo, Dongsheng LiICCV 2025 · 2 citations
