G^2VLM: Geometry Grounded Vision Language Model with Unified 3D Reconstruction and Spatial Reasoning
Wenbo hu, JINGLI LIN, Yilin Long, Yunlong Ran, Lihan Jiang, Yifan Wang, Chenming Zhu, Runsen Xu, Tai Wang, Jiangmiao Pang
摘要
Vision-Language Models (VLMs) still lack robustness in spatial intelligence, demonstrating poor performance on spatial understanding and reasoning tasks. We attribute this gap to the absence of a visual geometry learning process capable of reconstructing 3D space from 2D images. We present GVLM, a geometry grounded vision-language model that bridges two fundamental aspects of spatial intelligence: spatial 3D reconstruction and spatial understanding. GVLM natively leverages learned 3D visual geometry features to directly predict 3D attributes and enhance spatial reasoning tasks via in-context learning and interleaved reasoning. Our unified design is highly scalable for spatial understanding: it trains on abundant multi-view image and video data, while simultaneously leveraging the benefits of 3D visual priors that are typically only derived from hard-to-collect annotations.Experimental results demonstrate GVLM is proficient in both tasks, achieving comparable results to state-of-the-art feed-forward 3D reconstruction models and achieving better or competitive results across spatial understanding and reasoning tasks.By unifying a semantically strong VLM with low-level 3D vision tasks, we hope GVLM can serve as a strong baseline for the community and unlock more future applications, such as 3D scene editing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- From Correspondence to Actions: Human-Like Multi-Image Spatial Reasoning in Multi-modal Large Language ModelsMasanari Oi, Koki Maeda, Ryuto Koike, Daisuke Oba 等ICML 2026 · 被引用 2 次
- 3D-RFT: Reinforcement Fine-Tuning for Video-based 3D Scene UnderstandingXiongkun Linghu, Jiangyong Huang, Baoxiong Jia, Siyuan HuangICML 2026 · 被引用 1 次
它引用的顶会 Paper39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Sigmoid Loss for Language Image Pre-TrainingXiaohua Zhai, Basil Mustafa, Alexander Kolesnikov, Lucas BeyerICCV 2023 · 被引用 2,932 次
- Habitat 2.0: Training Home Assistants to Rearrange their HabitatAndrew Szot, Alexander Clegg, Eric Undersander, Erik Wijmans 等NeurIPS 2021 · 被引用 826 次
- Vision Transformers Need RegistersTimothée Darcet, Maxime Oquab, Julien Mairal, Piotr BojanowskiICLR 2024 · 被引用 769 次
相关 Paper
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 等CVPR 2026 · 被引用 171 次
- HiSpatial: Taming Hierarchical 3D Spatial Understanding in Vision-Language ModelsHuizhi Liang, Yichao Shen, Yu Deng, Sicheng Xu 等CVPR 2026 · 被引用 2 次
- Learning from Videos for 3D World: Enhancing MLLMs with 3D Vision Geometry PriorsDuo Zheng, Shijia Huang, Yanyang Li, Liwei WangNeurIPS 2025 · 被引用 130 次
- Proxy3D: Efficient 3D Representations for Vision-Language Models via Semantic Clustering and AlignmentJerry Jiang, Haowen Sun, Denis A. Gudovskiy, Yohei Nakata 等CVPR 2026 · 被引用 3 次
- Abstract 3D Perception for Spatial Intelligence in Vision-Language ModelsYifan Liu, Fangneng Zhan, Kaichen Zhou, Yilun Du 等CVPR 2026 · 被引用 6 次
