Dynamo-Depth: Fixing Unsupervised Depth Estimation for Dynamical Scenes
Yihong Sun, Bharath Hariharan
Abstract
Unsupervised monocular depth estimation techniques have demonstrated encouraging results but typically assume that the scene is static. These techniques suffer when trained on dynamical scenes, where apparent object motion can equally be explained by hypothesizing the object's independent motion, or by altering its depth. This ambiguity causes depth estimators to predict erroneous depth for moving objects. To resolve this issue, we introduce Dynamo-Depth, an unifying approach that disambiguates dynamical motion by jointly learning monocular depth, 3D independent flow field, and motion segmentation from unlabeled monocular videos. Specifically, we offer our key insight that a good initial estimation of motion segmentation is sufficient for jointly learning depth and independent motion despite the fundamental underlying ambiguity. Our proposed method achieves state-of-the-art performance on monocular depth estimation on Waymo Open [34] and nuScenes [3] Dataset with significant improvement in the depth of moving objects. Code and additional results are available at https://dynamo-depth.github.io .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ec3376e-8890-49fe-90ab-2c4495e4e8d4Cited by top-tier papers9
- ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy PredictionYi Feng, Yu Han, Xijing Zhang, Tanghui Li et al.AAAI 2025 · 8 citations
- Feed-Forward SceneDINO for Unsupervised Semantic Scene CompletionAleksandar Jevtic, Christoph Reich, Felix Wimbauer, Oliver Hahn et al.ICCV 2025 · 3 citations
- Self-Supervised Monocular 4D Scene Reconstruction for Egocentric VideosChengbo Yuan, Geng Chen, Li Yi, Yang GaoICCV 2025 · 2 citations
- EMD: Explicit Motion Modeling for High-Quality Street Gaussian SplattingXiaobao Wei, Qingpo Wuwu, Zhongyu Zhao, Zhuangzhe Wu et al.ICCV 2025 · 2 citations
- Scene-Centric Unsupervised Video Panoptic SegmentationChristoph Reich, Oliver Hahn, Nikita Araslanov, Laura Leal-Taixe et al.CVPR 2026 · 1 citation
Builds on17
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- Learning Monocular Depth in Dynamic Scenes via Instance-Aware Projection ConsistencySeokju Lee, Sunghoon Im, Stephen Lin, In So KweonAAAI 2021 · 107 citations
- RM-Depth: Unsupervised Learning of Recurrent Monocular Depth in Dynamic ScenesTak-Wai HuiCVPR 2022 · 62 citations
- Attentive and Contrastive Learning for Joint Depth and Motion Field EstimationSeokju Lee, François Rameau, Fei Pan, In So KweonICCV 2021 · 38 citations
Related papers
- AdaDepth: Exploiting Inherent Scene Information for Self-Supervised Depth Estimation in Dynamic ScenesXuanang Gao, Xiongbin Wu, Zhiwei Ning, Runze Yang et al.AAAI 2026
- Mining Supervision for Dynamic Regions in Self-Supervised Monocular Depth EstimationHoang Chuong Nguyen, Tianyu Wang, José M. Álvarez, Miaomiao LiuCVPR 2024 · 5 citations
- Multi-Object Discovery by Low-Dimensional Object MotionSadra Safadoust, Fatma GüneyICCV 2023 · 15 citations
- SynDeMo: Synergistic Deep Feature Alignment for Joint Learning of Depth and Ego-MotionBehzad Bozorgtabar, Mohammad Saeed Rad, Dwarikanath Mahapatra, Jean-Philippe ThiranICCV 2019 · 44 citations
- Instance-Level Video Depth in Groups Beyond OcclusionsYuan Liang, Yang Zhou, Ziming Sun, Tianyi Xiang et al.ICCV 2025
