SCE-SLAM: Scale-Consistent Monocular SLAM via Scene Coordinate Embeddings
Yuchen Wu, Jiahe Li, Xiaohan Yu, Lina Yu, Jin Zheng, Xiao Bai
摘要
Monocular visual SLAM enables 3D reconstruction from internet video and autonomous navigation on resource-constrained platforms, yet suffers from scale drift, i.e., the gradual divergence of estimated scale over long sequences. Existing frame-to-frame methods achieve real-time performance through local optimization but accumulate scale drift due to the lack of global constraints among independent windows. To address this, we propose SCE-SLAM, an end-to-end SLAM system that maintains scale consistency through scene coordinate embeddings, which are learned patch-level representations encoding 3D geometric relationships under a canonical scale reference. The framework consists of two key modules: geometry-guided aggregation that leverages 3D spatial proximity to propagate scale information from historical observations through geometry-modulated attention, and scene coordinate bundle adjustment that anchors current estimates to the reference scale through explicit 3D coordinate constraints decoded from the scene coordinate embeddings. Experiments on KITTI, Waymo, and vKITTI demonstrate substantial improvements: our method reduces absolute trajectory error by 8.36m on KITTI compared to the best prior approach, while maintaining 36 FPS and achieving scale consistency across large-scale scenes.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 被引用 1,248 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- iMAP: Implicit Mapping and Positioning in Real-TimeEdgar Sucar, Shikun Liu, Joseph Ortiz, Andrew J. DavisonICCV 2021 · 被引用 834 次
相关 Paper
- FoundationSLAM: Unleashing the Power of Depth Foundation Models for End-to-End Dense Visual SLAMYuchen Wu, Jiahe Li, Fabio Tosi, Matteo Poggi 等AAAI 2026
- MASt3R-SLAM: Real-Time Dense SLAM with 3D Reconstruction PriorsRiku Murai, Eric Dexheimer, Andrew J. DavisonCVPR 2025
- VGGT-Motion: Motion-Aware Calibration-Free Monocular SLAM for Long-Range ConsistencyZhuang Xiong, Chen Zhang, Qingshan Xu, Wenbing TaoICML 2026 · 被引用 6 次
- SCOPE: Scale-Consistent One-Pass Estimation of 3D GeometryZheng Zhang, Lihe Yang, Tianyu Yang, Chaohui Yu 等SIGGRAPH 2026
- Dynamic Visual SLAM using a General 3D PriorXingguang Zhong, Liren Jin, Marija Popovic, Jens Behley 等CVPR 2026 · 被引用 1 次
