VGGT: Visual Geometry Grounded Transformer
Jianyuan Wang, Minghao Chen, Nikita Karaev, Andrea Vedaldi, Christian Rupprecht, David Novotný
摘要
We present VGGT, a feed-forward neural network that directly infers all key 3D attributes of a scene, including camera parameters, point maps, depth maps, and 3D point tracks, from one, a few, or hundreds of its views. This approach is a step forward in 3D computer vision, where models have typically been constrained to and specialized for single tasks. It is also simple and efficient, reconstructing images in under one second, and still outperforming alternatives that require post-processing with visual geometry optimization techniques. The network achieves state-of-the-art results in multiple 3D tasks, including camera parameter estimation, multi-view depth estimation, dense point cloud reconstruction, and 3D point tracking. We also show that using pretrained VGGT as a feature backbone significantly enhances downstream tasks, such as non-rigid point tracking and feed-forward novel view synthesis. Code and models are publicly available at https://github.com/facebookresearch/vggt.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper517
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen 等ICLR 2026 · 被引用 720 次
- π3: Permutation-Equivariant Visual Geometry LearningYifan Wang, Jianjun Zhou, Haoyi Zhu, Wenzheng Chang 等ICLR 2026 · 被引用 318 次
- Spatial-MLLM: Boosting MLLM Capabilities in Visual-based Spatial IntelligenceDiankun Wu, Fangfu Liu, Yi-Hsin Hung, Yueqi DuanNeurIPS 2025 · 被引用 245 次
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldDominic Maggio, Hyungtae Lim, Luca CarloneNeurIPS 2025 · 被引用 176 次
- VLM-3R: Vision-Language Models Augmented with Instruction-Aligned 3D ReconstructionZhiwen Fan, Jian Zhang, Renjie Li, Junge Zhang 等CVPR 2026 · 被引用 171 次
它引用的顶会 Paper45
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun 等NeurIPS 2020 · 被引用 1,010 次
相关 Paper
- WorldMirror: Universal 3D World Reconstruction with Any-Prior PromptingYifan Liu, Zhiyuan Min, Zhenwei Wang, Junta Wu 等ICML 2026 · 被引用 58 次
- PAGE-4D: Disentangled Pose and Geometry Estimation for VGGT-4D PerceptionKaichen Zhou, Yuhan Wang, Grace Chen, Gaspard Beaudouin 等ICLR 2026 · 被引用 12 次
- Selfi: Self-improving Reconstruction Engine via 3D Geometric Feature AlignmentYouming Deng, Songyou Peng, Junyi Zhang, Kathryn Heal 等CVPR 2026 · 被引用 4 次
- GGPT: Geometry-Grounded Point TransformerYutong Chen, Yiming Wang, Xucong Zhang, Sergey Prokudin 等CVPR 2026 · 被引用 2 次
- V-DPM: 4D Video Reconstruction with Dynamic Point MapsEdgar Sucar, Eldar Insafutdinov, Zihang Lai, Andrea VedaldiCVPR 2026 · 被引用 29 次
