Deep Patch Visual Odometry
Zachary Teed, Lahav Lipson, Jia Deng
Abstract
We propose Deep Patch Visual Odometry (DPVO), a new deep learning system for monocular Visual Odometry (VO). DPVO uses a novel recurrent network architecture designed for tracking image patches across time. Recent approaches to VO have significantly improved the state-of-the-art accuracy by using deep networks to predict dense flow between video frames. However, using dense flow incurs a large computational cost, making these previous methods impractical for many use cases. Despite this, it has been assumed that dense flow is important as it provides additional redundancy against incorrect matches. DPVO disproves this assumption, showing that it is possible to get the best accuracy and efficiency by exploiting the advantages of sparse patch-based matching over dense flow. DPVO introduces a novel recurrent update operator for patch based correspondence coupled with differentiable bundle adjustment. On Standard benchmarks, DPVO outperforms all prior work, including the learning-based state-of-the-art VO-system (DROID) using a third of the memory while running 3x faster on average. Code is available at https: //github.com/princeton-vl/DPVO
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 492b86e5-ec31-48f0-8912-c18e666d6f5cCited by top-tier papers75
- WHAM: Reconstructing World-Grounded Humans with Accurate 3D MotionSoyong Shin, Juyong Kim, Eni Halilaj, Michael J. BlackCVPR 2024 · 66 citations
- OmniNav: A Unified Framework for Prospective Exploration and Visual-Language NavigationXinda Xue, Junjun Hu, Minghua Luo, Xie Shichao et al.ICLR 2026 · 51 citations
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 27 citations
- Scal3R: Scalable Test-Time Training for Large-Scale 3D ReconstructionTao Xie, Peishan Yang, Yudong Jin, Yingfeng Cai et al.CVPR 2026 · 26 citations
- Infinigen Indoors: Photorealistic Indoor Scenes using Procedural GenerationAlexander Raistrick, Lingjie Mei, Karhan Kayan, David Yan et al.CVPR 2024 · 24 citations
Builds on7
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Pixel-Perfect Structure-from-Motion with Featuremetric RefinementPhilipp Lindenberger, Paul-Edouard Sarlin, Viktor Larsson, Marc PollefeysICCV 2021 · 266 citations
- SuperGlue: Learning Feature Matching With Graph Neural NetworksPaul-Edouard Sarlin, Daniel DeTone, Tomasz Malisiewicz, Andrew RabinovichCVPR 2020
- LoFTR: Detector-Free Local Feature Matching With TransformersJiaming Sun, Zehong Shen, Yuang Wang, Hujun Bao et al.CVPR 2021
- Learning Accurate Dense Correspondences and When To Trust ThemPrune Truong, Martin Danelljan, Luc Van Gool, Radu TimofteCVPR 2021
Related papers
- Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual OdometryZhaoxing Zhang, Junda Cheng, Gangwei Xu, Xiaoxiang Wang et al.AAAI 2025 · 9 citations
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual OdometryNan Yang, Lukas von Stumberg, Rui Wang, Daniel CremersCVPR 2020
- Deep geometry-aware camera self-calibration from videoAnnika Hagemann, Moritz Knorr, Christoph StillerICCV 2023 · 33 citations
- Dual-Agent Reinforcement Learning for Adaptive and Cost-Aware Visual-Inertial OdometryFeiyang Pan, Shenghe Zheng, Chunyan Yin, Guangbin DouCVPR 2026
- From Variance to Veracity: Unbundling and Mitigating Gradient Variance in Differentiable Bundle Adjustment LayersSwaminathan Gurumurthy, Karnik Ram, Bingqing Chen, Zachary Manchester et al.CVPR 2024
