OpenVO: Open-World Visual Odometry with Temporal Dynamics Awareness
Phuc Nguyen, Anh N Nhu, Ming C. Lin
Abstract
We introduce OpenVO, a novel framework for Open-world Visual Odometry (VO) with temporal awareness under limited input conditions. OpenVO effectively estimates real-world–scale ego-motion from monocular dashcam footage with varying observation rates and uncalibrated cameras, enabling robust trajectory dataset construction from rare driving events recorded in dashcam.Existing VO methods are trained on fixed observation frequency (e.g., 10Hz or 12Hz), completely overlooking temporal dynamics information. Many prior methods also require calibrated cameras with known intrinsic parameters. Consequently, their performance degrades when (1) deployed under unseen observation frequencies or (2) applied to uncalibrated cameras. These significantly limit their generalizability to many downstream tasks, such as extracting trajectories from dashcam footage.To address these challenges, OpenVO (1) explicitly encodes temporal dynamics information within a two-frame pose regression framework and (2) leverages 3D geometric priors derived from foundation models. We validate our method on three major autonomous-driving benchmarks – KITTI, nuScenes, and Argoverse 2 – achieving more than 20% performance improvement over state-of-the-art approaches. Under varying observation rate settings, our method is significantly more robust, achieving 46%–92% lower errors across all metrics.These results demonstrate the versatility of OpenVO for real-world 3D reconstruction and diverse downstream applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 618496ef-30c4-4e3d-95b7-28b3c18c002cBuilds on26
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Metric3D: Towards Zero-shot Metric 3D Prediction from A Single ImageWei Yin, Chi Zhang, Hao Chen, Zhipeng Cai et al.ICCV 2023 · 388 citations
- VectorMapNet: End-to-end Vectorized HD Map LearningYicheng Liu, Tianyuan Yuan, Yue Wang, Yilun Wang et al.ICML 2023 · 332 citations
Related papers
- XVO: Generalized Visual Odometry via Cross-Modal Self-TrainingLei Lai, Zhongkai Shangguan, Jimuyang Zhang, Eshed Ohn-BarICCV 2023 · 27 citations
- ZeroVO: Visual Odometry with Minimal AssumptionsLei Lai, Zekai Yin, Eshed Ohn-BarCVPR 2025
- StreamVLO: Streaming Visual-LiDAR Odometry with Cumulative Drift CompensationMengmeng Liu, Jiuming Liu, Michael Ying Yang, Chaokang Jiang et al.CVPR 2026
- Joint Monocular 3D Vehicle Detection and TrackingHou-Ning Hu, Qi-Zhi Cai, Dequan Wang, Ji Lin et al.ICCV 2019 · 242 citations
- R3D3: Dense 3D Reconstruction of Dynamic Scenes from Multiple CamerasAron Schmied, Tobias Fischer, Martin Danelljan, Marc Pollefeys et al.ICCV 2023 · 52 citations
