Video Spatial Reasoning with Object-Centric 3D Rollout
Haoran Tang, Meng Cao, Ruyang Liu, Xiaoxi Liang, Linglong Li, Ge Li, Xiaodan Liang
Abstract
Video-R1: (Only mentioned objects) <think>Let me think about this carefully. If I am standing by the pillow and facing the radiator, I need to determine the direction of the laptop relative to my position. Wait, let's break it down. If I am facing the radiator, that means I am facing towards the back of the room. The laptop is on the right side of the room, so it would be to my right. Hmm, if I am facing the radiator and the laptop is to my right, it would be to my front-right. Oh, I see. The laptop is not to my back, so options C and D are not correct. The laptop is not to my left, so option A is not correct. Therefore, the laptop is to my front-right.</think> <answer>B</answer> Ours: (More relevant objects for spatial reasoning) <think>Let me think about this for a moment. If I am standing by the pillow and facing the radiator, I need to determine the position of the laptop relative to me. First, let's break it down. If I am facing the radiator, I am looking towards the back of the room. The pillow is to my left, and it is to the right of the bed, so if I am standing by the pillow, I am facing the radiator from the front. Now, let's consider the laptop. The laptop is on the desk, and the desk is to the left of the bed. If I am facing the radiator, the laptop is not directly in front of me, but rather to the left of the bed. Hmm, if I am standing by the pillow and facing the radiator, the laptop is not to my back, as it is to the left of the bed. It is also not to my right, as it is to the left of the bed. Wait, if the laptop is to the left of the bed and I am facing the radiator, it must be to my front-left. Oh, I see. The laptop is to my front-left. Let's verify. If I am standing by the pillow and facing the radiator, the laptop is indeed to my front-left. Therefore, the laptop is to my front-left.</think> <answer>A</answer> Problem: Object Relative Direction Question: If I am standing by the pillow and facing the radiator, is the laptop to my A. front-left, B. front-right, C. back-left, or D. back-right? (Ground Truth: A) ries. Experiments demonstrate state-of-the-art performance: our 3B-parameter model achieves 47.5% accuracy on VSI-Bench, outperforming several 7B baselines. Ablations confirm OCR's superiority over prior rollout strategies (e.g., T-GRPO, NoisyRollout).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- ReVSI: Rebuilding Visual Spatial Intelligence Evaluation for Accurate Assessment of VLM 3D ReasoningYiming Zhang, Jiacheng Chen, Jiaqi Tan, Yongsen Mao et al.ICML 2026 · 7 citations
- VideoKR: Towards Knowledge- and Reasoning-Intensive Video UnderstandingLin Fu, Zheyuan Yang, Yang Wang, Tingyu Song et al.ICML 2026 · 1 citation
Builds on13
- ScanNet++: A High-Fidelity Dataset of 3D Indoor ScenesChandan Yeshwanth, Yueh-Cheng Liu, Matthias Nießner, Angela DaiICCV 2023 · 659 citations
- SpatialRGPT: Grounded Spatial Reasoning in Vision-Language ModelsAn-Chieh Cheng, Hongxu Yin, Yang Fu, Qiushan Guo et al.NeurIPS 2024 · 412 citations
- DeepEyes: Incentivizing "Thinking with Images" via Reinforcement LearningZiwei Zheng, Michael Yang, Jack Hong, Chenxiao Zhao et al.ICLR 2026 · 321 citations
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 245 citations
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang et al.CVPR 2026 · 171 citations
Related papers
- PerLA: Perceptive 3D Language AssistantGuofeng Mei, Wei Lin, Luigi Riz, Yujiao Wu et al.CVPR 2025
- FoREST: Frame of Reference Evaluation in Spatial Reasoning TasksTanawan Premsri, Parisa KordjamshidiEMNLP 2025
- SeeGround: See and Ground for Zero-Shot Open-Vocabulary 3D Visual GroundingRong Li, Shijie Li, Lingdong Kong, Xulei Yang et al.CVPR 2025
- Action-Sketcher: From Reasoning to Action via Visual Sketches for Robotic ManipulationHuajie Tan, Peterson Co, Yijie Xu, Shanyu Rong et al.CVPR 2026
- Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference under AmbiguitiesZheyuan Zhang, Fengyuan Hu, Jayjun Lee, Freda Shi et al.ICLR 2025
