SpatialLadder: Progressive Training for Spatial Reasoning in Vision-Language Models
Hongxing Li, Dingming Li, Zixuan Wang, Yuchen Yan, Hang Wu, Wenqi Zhang, Yongliang Shen, Weiming Lu, Jun Xiao, Yueting Zhuang
Abstract
Spatial reasoning remains a fundamental challenge for Vision-Language Models (VLMs), with current approaches struggling to achieve robust performance despite recent advances. We identify that this limitation stems from a critical gap: existing methods attempt to learn spatial reasoning directly without establishing the hierarchical foundations of perception and understanding. To address this challenge, we present a comprehensive methodology for building spatial intelligence progressively. We introduce SpatialLadder-26k, a multimodal dataset containing 26,610 samples spanning object localization, single-image, multi-view, and video spatial reasoning tasks, constructed through a standardized pipeline that ensures systematic coverage across modalities. Building on this dataset, we design a three-stage progressive training framework that (1) establishes spatial perception through object localization, (2) develops spatial understanding through multi-dimensional spatial tasks, and (3) strengthens complex reasoning via reinforcement learning with verifiable rewards. This approach yields SpatialLadder, a 3B-parameter model that achieves state-of-the-art performance on spatial reasoning benchmarks, with 23.4% average improvement over the base model, surpassing GPT-4o by 20.8% and Gemini-2.0-Flash by 10.1%. Notably, SpatialLadder maintains strong generalization with 7.2% improvement on out-of-domain benchmarks, demonstrating that progressive training from perception to reasoning is essential for robust spatial intelligence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5cb4e6c0-a346-4151-9ea3-19d5981ad935Cited by top-tier papers15
- Scaling Spatial Intelligence with Multimodal Foundation ModelsZhongang Cai, Wang Ruisi, Chenyang Gu, Fanyi Pu et al.CVPR 2026 · 81 citations
- Think with 3D: Geometric Imagination Grounded Spatial Reasoning from Limited ViewsZhangquan Chen, Manyuan Zhang, Xinlei Yu, Xufang Luo et al.CVPR 2026 · 61 citations
- Spatial-SSRL: Enhancing Spatial Understanding via Self-Supervised Reinforcement LearningYuhong Liu, Beichen Zhang, Yuhang Zang, Yuhang Cao et al.CVPR 2026 · 43 citations
- Geometrically-Constrained Agent for Spatial ReasoningZeren Chen, Xiaoya Lu, Zhijie Zheng, Pengrui Li et al.CVPR 2026 · 29 citations
- SpaceMind: Camera-Guided Modality Fusion for Spatial Reasoning in Vision-Language ModelsRuosen Zhao, Zhikang Zhang, Jialei Xu, Jiahao Chang et al.CVPR 2026 · 21 citations
Builds on20
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMsPeter Tong, Ellis Brown, Penghao Wu, Sanghyun Woo et al.NeurIPS 2024 · 1,004 citations
- Vision-R1: Incentivizing Reasoning Capability in Multimodal Large Language ModelsWenxuan Huang, Bohan Jia, Shaosheng Cao, Zheyu Ye et al.ICLR 2026 · 670 citations
- 3D-LLM: Injecting the 3D World into Large Language ModelsYining Hong, Haoyu Zhen, Peihao Chen, Shuhong Zheng et al.NeurIPS 2023 · 662 citations
Related papers
- HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language ModelsHuizhi Liang, Yichao Shen, Yu Deng, Sicheng Xu et al.CVPR 2026 · 2 citations
- Aligning Cross-View Visual Geometries in LVLMs Through Human-Like Reasoning LearningYuming Qiao, Liang Luo, Dan Meng, Yifan Yang et al.AAAI 2026
- Vision-G1: Towards General Reasoning Vision-Language Models via Reinforcement LearningYuheng Zha, Kun Zhou, Yujia Wu, Yushu Wang et al.AAAI 2026
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
- SpatialReasoner: Towards Explicit and Generalizable 3D Spatial ReasoningWufei Ma, Yu-Cheng Chou, Qihao Liu, Xingrui Wang et al.NeurIPS 2025 · 77 citations
