VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) Manifold
Dominic Maggio, Hyungtae Lim, Luca Carlone
摘要
We present VGGT-SLAM, a dense RGB SLAM system constructed by incrementally and globally aligning submaps created from the feed-forward scene reconstruction approach VGGT using only uncalibrated monocular cameras. While related works align submaps using similarity transforms (i.e., translation, rotation, and scale), we show that such approaches are inadequate in the case of uncalibrated cameras. In particular, we revisit the idea of reconstruction ambiguity, where given a set of uncalibrated cameras with no assumption on the camera motion or scene structure, the scene can only be reconstructed up to a 15-degreesof-freedom projective transformation of the true geometry. This inspires us to recover a consistent scene reconstruction across submaps by optimizing over the SL(4) manifold, thus estimating 15-degrees-of-freedom homography transforms between sequential submaps while accounting for potential loop closure constraints. As verified by extensive experiments, we demonstrate that VGGT-SLAM achieves improved map quality using long video sequences that are infeasible for VGGT due to its high GPU requirements. Our code is available at https://github.com/MIT-SPARK/VGGT-SLAM.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper30
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen 等ICLR 2026 · 被引用 720 次
- TTT3R: 3D Reconstruction as Test-Time TrainingXingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger 等ICLR 2026 · 被引用 139 次
- Geometry Forcing: Marrying Video Diffusion and 3D Representation for Consistent World ModelingHaoyu Wu, Diankun Wu, Tianyu He, Junliang Guo 等ICLR 2026 · 被引用 89 次
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai 等CVPR 2026 · 被引用 26 次
- OmniVGGT: Omni-Modality Driven Visual Geometry Grounded TransformerHaosong Peng, Hao Li, Yalun Dai, Yushi Lan 等CVPR 2026 · 被引用 22 次
它引用的顶会 Paper14
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Gaussian Splatting SLAMHidenobu Matsuki, Riku Murai, Paul H. J. Kelly, Andrew J. DavisonCVPR 2024 · 被引用 328 次
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
相关 Paper
- VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range ConsistencyZhuang Xiong, Chen Zhang, Qingshan Xu, Wenbing TaoICML 2026 · 被引用 6 次
- SLAM3R: Real-Time Dense Scene Reconstruction from Monocular RGB VideosYuzheng Liu, Siyan Dong, Shuzhe Wang, Yingda Yin 等CVPR 2025
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsRiku Murai, Eric Dexheimer, Andrew J. DavisonCVPR 2025
- Unblur-SLAM: Dense Neural SLAM for Blurry InputsQi Zhang, Denis Rozumny, Francesco Girlanda, Sezer Karaoglu 等CVPR 2026 · 被引用 1 次
- AMB3R: Accurate Feed-forward Metric-scale 3D Reconstruction with BackendHengyi Wang, Lourdes AgapitoCVPR 2026 · 被引用 17 次
