VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range Consistency
Zhuang Xiong, Chen Zhang, Qingshan Xu, Wenbing Tao
Abstract
Despite recent progress in calibration-free monocular SLAM via 3D vision foundation models, scale drift remains severe on long sequences. Motion-agnostic partitioning breaks contextual coherence and causes zero-motion drift, while conventional geometric alignment is computationally expensive. To address these issues, we propose VGGT-Motion, a calibration-free SLAM system for efficient and robust global consistency over kilometer-scale trajectories. Specifically, we first propose a motion-aware submap construction mechanism that uses optical flow to guide adaptive partitioning, prune static redundancy, and encapsulate turns for stable local geometry. We then design an anchor-driven direct Sim(3) registration strategy. By exploiting context-balanced anchors, it achieves search-free, pixel-wise dense alignment and efficient loop closure without costly feature matching. Finally, a lightweight submap-level pose graph optimization enforces global consistency with linear complexity, enabling scalable long-range operation. Experiments show that VGGT-Motion markedly improves trajectory accuracy and efficiency, achieving state-of-the-art performance in zero-shot, long-range calibration-free monocular SLAM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c382e26-37c6-4398-8e55-6aa2efd9ab3eBuilds on14
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
Related papers
- VGGT-SLAM: Dense RGB SLAM Optimized on the SL(4) ManifoldDominic Maggio, Hyungtae Lim, Luca CarloneNeurIPS 2025 · 176 citations
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsRiku Murai, Eric Dexheimer, Andrew J. DavisonCVPR 2025
- FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAMYuchen Wu, Jiahe Li, Fabio Tosi, Matteo Poggi et al.AAAI 2026
- TALO: Pushing 3D Vision Foundation Models Towards Globally Consistent Online ReconstructionFengyi Zhang, Tianjun Zhang, Kasra Khosoussi, Zheng Zhang et al.CVPR 2026 · 4 citations
- SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate EmbeddingsYuchen Wu, Jiahe Li, Xiaohan Yu, Lina Yu et al.CVPR 2026 · 1 citation
