3ViewSense: Spatial and Mental Perspective Reasoning from Orthographic Views in Vision-Language Models
Shaoxiong Zhan, Yanlin Lai, Zheng Liu, Zijian Lin, Lin Hai, Xiaodong Cai, Shen Li, Wen Huang, Hai-Tao Zheng
摘要
Current Large Language Models have achieved Olympiad-level logic, yet Vision-Language Models paradoxically falter on elementary spatial tasks like block counting. This capability mismatch reveals a critical "spatial intelligence gap," where models fail to construct coherent 3D mental representations from 2D observations. We uncover this gap via diagnostic analyses showing the bottleneck is a missing view-consistent spatial interface rather than insufficient visual features or weak reasoning. To bridge this, we introduce 3ViewSense, a framework that grounds spatial reasoning in Orthographic Views. Drawing on engineering cognition, we propose a "Simulate-and-Reason" mechanism that decomposes complex scenes into canonical orthographic projections to resolve geometric ambiguities. By aligning egocentric perceptions with these allocentric references, our method facilitates explicit mental rotation and reconstruction. Empirical results on spatial reasoning benchmarks demonstrate that our method significantly outperforms existing baselines, with consistent gains on occlusion-heavy counting and view-consistent spatial reasoning. The framework also improves the stability and consistency of spatial descriptions, offering a scalable path toward stronger spatial intelligence in multimodal systems. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper11
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 被引用 245 次
- RL's Razor: Why Online Reinforcement Learning Forgets LessIdan Shenfeld, Jyothish Pari, Pulkit AgrawalICLR 2026 · 被引用 176 次
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 等CVPR 2026 · 被引用 171 次
- OmniSpatial: Towards Comprehensive Spatial Reasoning Benchmark for Vision Language ModelsMengdi Jia, Zekun Qi, Shaochen Zhang, Wenyao Zhang 等ICLR 2026 · 被引用 109 次
- Spatial Mental Modeling from Limited ViewsQineng Wang, Baiqiao Yin, Pingyue Zhang, Jianshu Zhang 等ICLR 2026 · 被引用 92 次
相关 Paper
- Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and AlignmentJerry Jiang, Haowen Sun, Denis A. Gudovskiy, Yohei Nakata 等CVPR 2026 · 被引用 3 次
- Pursuing Minimal Sufficiency in Spatial ReasoningYejie Guo, Yunzhong Hou, Wufei Ma, Meng Tang 等ICLR 2026 · 被引用 3 次
- G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial ReasoningWenbo hu, JINGLI LIN, Yilin Long, Yunlong Ran 等CVPR 2026
- Keep it SymPL: Symbolic Projective Layout for Allocentric Spatial Reasoning in Vision-Language ModelsJaeyun Jang, Seunghui Shin, Taeho Park, Hyoseok HwangCVPR 2026 · 被引用 1 次
- SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical EvaluationWenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang 等ACL 2025 · 被引用 20 次
