STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li, Yu Rong, Tingyang Xu, Ziqi Zhang, Deli Zhao, Wenbing Huang
Abstract
Multimodal Large Language Models (MLLMs) remain far from human-level performance in multi-view spatial reasoning, where models must establish object correspondences across view and infer coherent scene semantics. We analyze this limitation through the Transformation-Driven Visual Reasoning (TVR) task and find that Supervised Fine-Tuning (SFT) fails to capture cross-view consistency, whereas reinforcement learning (RL) fails to reliably identify key referential objects. To bridge this gap, we introduce multi-View Spatial TrAnsformation Reasoning (STAR-R1), a two-stage framework that combines process-supervised SFT with a referential-aware RL paradigm. STAR-R1 first learns structured spatial reasoning trajectories from high-quality CoTs and then uses fine-grained rewards on referential selection and answer correctness to encourage effective exploration and robust scene interpretation. Despite using only a small amount of high-quality training data, STAR-R1 surpasses state-of-the-art models with far more training data on the multi-view spatial understanding benchmarks TVR, MMSI-Bench, MindCube-Bench, and SPAR-Bench. Our study reveals the overlooked potential of RL in multi-view spatial understanding and points a way toward potentially achieving more human-like spatial reasoning in MLLMs.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 28cedca7-3b71-49bd-8bd8-b16a3901047bBuilds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo et al.NeurIPS 2025 · 528 citations
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang et al.ICLR 2026 · 195 citations
- Towards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningWenkai Yang, Shuming Ma, Yankai Lin, Furu WeiNeurIPS 2025 · 141 citations
Related papers
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language ModelsXiaoyu Zhan, Wenxuan Huang, Hao Sun, Xinyu Fu et al.NeurIPS 2025 · 11 citations
- The Art of Interrogation: Consistency Amplifies Factuality in Spatial ReasoningThéo Uscidda, Marta Gazulla, Maks Ovsjanikov, Federico Tombari et al.ICML 2026
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline MatchingHao Zhong, Muzhi Zhu, Shenyan Zeng, Anzhou Li et al.CVPR 2026 · 1 citation
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesYihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng et al.NeurIPS 2025 · 61 citations
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao et al.NeurIPS 2025 · 41 citations
