STAR-R1: Multi-View Spatial TrAnsformation Reasoning by Reinforcing Multimodal LLMs
Zongzhao Li, Zongyang Ma, Mingze Li, Songyou Li, Yu Rong, Tingyang Xu, Ziqi Zhang, Deli Zhao, Wenbing Huang
摘要
Multimodal Large Language Models (MLLMs) remain far from human-level performance in multi-view spatial reasoning, where models must establish object correspondences across view and infer coherent scene semantics. We analyze this limitation through the Transformation-Driven Visual Reasoning (TVR) task and find that Supervised Fine-Tuning (SFT) fails to capture cross-view consistency, whereas reinforcement learning (RL) fails to reliably identify key referential objects. To bridge this gap, we introduce multi-View Spatial TrAnsformation Reasoning (STAR-R1), a two-stage framework that combines process-supervised SFT with a referential-aware RL paradigm. STAR-R1 first learns structured spatial reasoning trajectories from high-quality CoTs and then uses fine-grained rewards on referential selection and answer correctness to encourage effective exploration and robust scene interpretation. Despite using only a small amount of high-quality training data, STAR-R1 surpasses state-of-the-art models with far more training data on the multi-view spatial understanding benchmarks TVR, MMSI-Bench, MindCube-Bench, and SPAR-Bench. Our study reveals the overlooked potential of RL in multi-view spatial understanding and points a way toward potentially achieving more human-like spatial reasoning in MLLMs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 被引用 11,349 次
- Video-R1: Reinforcing Video Reasoning in MLLMsKaituo Feng, Kaixiong Gong, Bohao Li, Zonghao Guo 等NeurIPS 2025 · 被引用 528 次
- MMSI-Bench: A Benchmark for Multi-Image Spatial IntelligenceSihan Yang, Runsen Xu, Yiman Xie, Sizhe Yang 等ICLR 2026 · 被引用 195 次
- Towards Thinking-Optimal Scaling of Test-Time Compute for LLM ReasoningWenkai Yang, Shuming Ma, Yankai Lin, Furu WeiNeurIPS 2025 · 被引用 141 次
相关 Paper
- Actial: Activate Spatial Reasoning Ability of Multimodal Large Language ModelsXiaoyu Zhan, Wenxuan Huang, Hao Sun, Xinyu Fu 等NeurIPS 2025 · 被引用 11 次
- The Art of Interrogation: Consistency Amplifies Factuality in Spatial ReasoningThéo Uscidda, Marta Gazulla, Maks Ovsjanikov, Federico Tombari 等ICML 2026
- Eliciting Complex Spatial Reasoning in MLLMs through Wide-Baseline MatchingHao Zhong, Muzhi Zhu, Shenyan Zeng, Anzhou Li 等CVPR 2026 · 被引用 1 次
- OpenVLThinker: Complex Vision-Language Reasoning via Iterative SFT-RL CyclesYihe Deng, Hritik Bansal, Fan Yin, Nanyun Peng 等NeurIPS 2025 · 被引用 61 次
- Fine-Grained Preference Optimization Improves Spatial Reasoning in VLMsYifan Shen, Yuanzhe Liu, Jingyuan Zhu, Xu Cao 等NeurIPS 2025 · 被引用 41 次
