DeepV2D: Video to Depth with Differentiable Structure from Motion
Zachary Teed, Jia Deng
Abstract
We propose DeepV2D, an end-to-end deep learning architecture for predicting depth from video. DeepV2D combines the representation ability of neural networks with the geometric principles governing image formation. We compose a collection of classical geometric algorithms, which are converted into trainable modules and combined into an end-to-end differentiable architecture. DeepV2D interleaves two stages: motion estimation and depth estimation. During inference, motion and depth estimation are alternated and converge to accurate depth. Code is available this https URL.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 372ab969-34c4-441c-a1d3-a8dd768a4812Cited by top-tier papers97
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth Anything 3: Recovering the Visual Space from Any ViewsHaotong Lin, Sili Chen, Jun Hao Liew, Donny Y. Chen et al.ICLR 2026 · 720 citations
- NICE-SLAM: Neural Implicit Scalable Encoding for SLAMZihan Zhu, Songyou Peng, Viktor Larsson, Weiwei Xu et al.CVPR 2022 · 720 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii et al.CVPR 2024 · 302 citations
Related papers
- UprightNet: Geometry-Aware Camera Orientation Estimation From Single ImagesWenqi Xian, Zhengqi Li, Noah Snavely, Matthew Fisher et al.ICCV 2019 · 52 citations
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 207 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- DSGN: Deep Stereo Geometry Network for 3D Object DetectionYilun Chen, Shu Liu, Xiaoyong Shen, Jiaya JiaCVPR 2020
- SpatialTrackerV2: Advancing 3D Point Tracking with Explicit Camera MotionYuxi Xiao, Jianyuan Wang, Nan Xue, Nikita Karaev et al.ICCV 2025 · 6 citations
