Leveraging Consistent Spatio-Temporal Correspondence for Robust Visual Odometry
Zhaoxing Zhang, Junda Cheng, Gangwei Xu, Xiaoxiang Wang, Can Zhang, Xin Yang
Abstract
Recent approaches to VO have significantly improved performance by using deep networks to predict optical flow between video frames. However, existing methods still suffer from noisy and inconsistent flow matching, making it difficult to handle challenging scenarios and long-sequence estimation. To overcome these challenges, we introduce Spatio-Temporal Visual Odometry (STVO), a novel deep network architecture that effectively leverages inherent spatiotemporal cues to enhance the accuracy and consistency of multi-frame flow matching. With more accurate and consistent flow matching, STVO can achieve better pose estimation through the bundle adjustment (BA). Specifically, STVO introduces two innovative components: 1) the Temporal Propagation Module that utilizes multi-frame information to extract and propagate temporal cues across adjacent frames, maintaining temporal consistency; 2) the Spatial Activation Module that utilizes geometric priors from the depth maps to enhance spatial consistency while filtering out excessive noise and incorrect matches. Our STVO achieves state-ofthe-art performance on TUM-RGBD, EuRoc MAV, ETH3D and KITTI Odometry benchmarks. Notably, it improves accuracy by 77.8% on ETH3D benchmark and 38.9% on KITTI Odometry benchmark over the previous best methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on11
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Causal Intervention for Weakly-Supervised Semantic SegmentationDong Zhang, Hanwang Zhang, Jinhui Tang, Xian-Sheng Hua et al.NeurIPS 2020 · 563 citations
- Learning to Estimate Hidden Motions with Global Motion AggregationShihao Jiang, Dylan Campbell, Yao Lu, Hongdong Li et al.ICCV 2021 · 402 citations
- Deep Patch Visual OdometryZachary Teed, Lahav Lipson, Jia DengNeurIPS 2023 · 323 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
Related papers
- Generalizing to the Open World: Deep Visual Odometry With Online AdaptationShunkai Li, Xin Wu, Yingdian Cao, Hongbin ZhaCVPR 2021
- Joint Spatial-Temporal Optimization for Stereo 3D Object TrackingPeiliang Li, Jieqi Shi, Shaojie ShenCVPR 2020
- D3VO: Deep Depth, Deep Pose and Deep Uncertainty for Monocular Visual OdometryNan Yang, Lukas von Stumberg, Rui Wang, Daniel CremersCVPR 2020
- Deep Two-View Structure-From-Motion RevisitedJianyuan Wang, Yiran Zhong, Yuchao Dai, Stan Birchfield et al.CVPR 2021
- ARFlow: Auto-regressive Optical Flow Estimation for Arbitrary-Length Videos via Progressive Next-Frame ForecastingJiuming Liu, Mengmeng Liu, Siting Zhu, Yunpeng Zhang et al.ICLR 2026
