LightedDepth: Video Depth Estimation in Light of Limited Inference View Angles
Shengjie Zhu, Xiaoming Liu
Abstract
Video depth estimation infers the dense scene depth from immediate neighboring video frames. While recent works consider it a simplified structure-from-motion (SfM) problem, it still differs from the SfM in that significantly fewer view angels are available in inference. This setting, however, suits the mono-depth and optical flow estimation. This observation motivates us to decouple the video depth estimation into two components, a normalized pose estimation over a flowmap and a logged residual depth estimation over a mono-depth map. The two parts are unified with an efficient off-the-shelf scale alignment algorithm. Additionally, we stabilize the indoor two-view pose estimation by including additional projection constraints and ensuring sufficient camera translation. Though a two-view algorithm, we validate the benefit of the decoupling with the substantial performance improvement over multi-view iterative prior works on indoor and outdoor datasets. Codes and models are available at https://github.com/ShngJZ/LightedDepth .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Tame a Wild Camera: In-the-Wild Monocular Camera CalibrationShengjie Zhu, Abhinav Kumar, Masa Hu, Xiaoming LiuNeurIPS 2023 · 47 citations
- PTC-Depth: Pose-Refined Monocular Depth Estimation with Temporal ConsistencyLeezy Han, Seunggyu Kim, Dongseok Shim, Hyeonbeom LeeCVPR 2026
- Vision-Language Embodiment for Monocular Depth EstimationJinchang Zhang, Guoyu LuCVPR 2025
Builds on12
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- DROID-SLAM: Deep Visual SLAM for Monocular, Stereo, and RGB-D CamerasZachary Teed, Jia DengNeurIPS 2021 · 1,248 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- DeepV2D: Video to Depth with Differentiable Structure from MotionZachary Teed, Jia DengICLR 2020 · 314 citations
- Multi-View Depth Estimation by Fusing Single-View Depth Probability with Multi-View GeometryGwangbin Bae, Ignas Budvytis, Roberto CipollaCVPR 2022 · 58 citations
Related papers
- Deep Two-View Structure-From-Motion RevisitedJianyuan Wang, Yiran Zhong, Yuchao Dai, Stan Birchfield et al.CVPR 2021
- Scale-flow: Estimating 3D Motion from VideoHan Ling, Quansen Sun, Zhenwen Ren, Yazhou Liu et al.ACM MM 2022 · 7 citations
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 44 citations
- Can Scale-Consistent Monocular Depth Be Learned in a Self-Supervised Scale-Invariant Manner?Lijun Wang, Yifan Wang, Linzhao Wang, Yunlong Zhan et al.ICCV 2021 · 48 citations
- RePoseD: Efficient Relative Pose Estimation With Known Depth InformationYaqing Ding, Viktor Kocur, Václav Vávra, Zuzana Berger Haladová et al.ICCV 2025 · 2 citations
