MambaVO: Deep Visual Odometry Based on Sequential Matching Refinement and Training Smoothing
Shuo Wang, Wanting Li, Yongcai Wang, Zhaoxin Fan, Zhe Huang, Xudong Cai, Jian Zhao, Deying Li
Abstract
Deep visual odometry has demonstrated great advancements by learning-to-optimize technology. This approach heavily relies on the visual matching across frames. However, ambiguous matching in challenging scenarios leads to significant errors in geometric modeling and bundle adjustment optimization, which undermines the accuracy and robustness of pose estimation. To address this challenge, this paper proposes MambaVO, which conducts robust initialization, Mamba-based sequential matching refinement, and smoothed training to enhance the matching quality and improve the pose estimation. Specifically, the new frame is matched with the closest keyframe in the maintained Point-Frame Graph (PFG) via the semi-dense based Geometric Initialization Module (GIM). Then the initialized PFG is processed by a proposed Geometric Mamba Module (GMM), which exploits the matching features to refine the overall inter-frame matching. The refined PFG is finally processed by differentiable BA to optimize the poses and the map. To deal with the gradient variance, a Trending-Aware Penalty (TAP) is proposed to smooth training and enhance convergence and stability. A loop closure module is finally applied to enable MambaVO++. On public benchmarks, MambaVO and MambaVO++ demonstrate SOTA performance, while ensuring real-time running.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5093f459-66df-440f-8c54-1d5a2878f3cdCited by top-tier papers4
- Aux-Think: Exploring Reasoning Strategies for Data-Efficient Vision-Language NavigationShuo Wang, Yongcai Wang, Wanting Li, Xudong Cai et al.NeurIPS 2025 · 28 citations
- Progress-Think: Semantic Progress Reasoning for Vision-Language NavigationShuo Wang, Yucheng Wang, Guoxin Lian, Yongcai Wang et al.CVPR 2026 · 10 citations
- DNOI-4DRO: Deep 4D Radar Odometry with Differentiable Neural-Optimization IterationsShouyi Lu, Huanyu Zhou, Guirong Zhuo, Xiao TangAAAI 2026 · 2 citations
- StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift CompensationMengmeng Liu, Jiuming Liu, Michael Ying Yang, Chaokang Jiang et al.CVPR 2026
Builds on18
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Efficiently Modeling Long Sequences with Structured State SpacesAlbert Gu, Karan Goel, Christopher RéICLR 2022 · 3,482 citations
- Vision Mamba: Efficient Visual Representation Learning with Bidirectional State Space ModelLianghui Zhu, Bencheng Liao, Qian Zhang, Xinlong Wang et al.ICML 2024 · 1,725 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- On the Parameterization and Initialization of Diagonal State Space ModelsAlbert Gu, Karan Goel, Ankit Gupta, Christopher RéNeurIPS 2022 · 690 citations
Related papers
- Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual OdometryZhaoxing Zhang, Junda Cheng, Gangwei Xu, Xiaoxiang Wang et al.AAAI 2025 · 9 citations
- Generalizing to the Open World: Deep Visual Odometry With Online AdaptationShunkai Li, Xin Wu, Yingdian Cao, Hongbin ZhaCVPR 2021
- From Variance to Veracity: Unbundling and Mitigating Gradient Variance in Differentiable Bundle Adjustment LayersSwaminathan Gurumurthy, Karnik Ram, Bingqing Chen, Zachary Manchester et al.CVPR 2024
- GO-SLAM: Global Optimization for Consistent 3D Instant ReconstructionYoumin Zhang, Fabio Tosi, Stefano Mattoccia, Matteo PoggiICCV 2023 · 208 citations
- PS-Mamba: Spatial-Temporal Graph Mamba for Pose Sequence RefinementHaoye Dong, Gim Hee LeeICCV 2025
