SfM-TTR: Using Structure from Motion for Test-Time Refinement of Single-View Depth Networks
Sergio Izquierdo, Javier Civera
Abstract
Estimating a dense depth map from a single view is geometrically ill-posed, and state-of-the-art methods rely on learning depth's relation with visual appearance using deep neural networks. On the other hand, Structure from Motion (SfM) leverages multi-view constraints to produce very accurate but sparse maps, as matching across images is typically limited by locally discriminative texture. In this work, we combine the strengths of both approaches by proposing a novel test-time refinement (TTR) method, denoted as SfM-TTR, that boosts the performance of single-view depth networks at test time using SfM multi-view cues. Specifically, and differently from the state of the art, we use sparse SfM point clouds as test-time self-supervisory signal, fine-tuning the network encoder to learn a better representation of the test scene. Our results show how the addition of SfM-TTR to several state-of-the-art self-supervised and supervised networks improves significantly their performance, outperforming previous TTR baselines mainly based on photometric multi-view consistency. The code is available at https://github.com/serizba/SfM-TTR .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- LightDepth: Single-View Depth Self-Supervision from Illumination DeclineJavier Rodriguez Puigvert, Victor M. Batlle, J. M. M. Montiel, Ruben Martinez-Cantin et al.ICCV 2023 · 22 citations
- MVSAnywhere: Zero-Shot Multi-View StereoSergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando et al.CVPR 2025
- Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric VideosVadim Tschernezki, Diane Larlus, Iro Laina, Andrea VedaldiCVPR 2025
Builds on12
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu et al.CVPR 2022 · 320 citations
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 265 citations
Related papers
- Test3R: Learning to Reconstruct 3D at Test TimeYuheng Yuan, Qiuhong Shen, Shizun Wang, Xingyi Yang et al.NeurIPS 2025 · 19 citations
- The Temporal Opportunist: Self-Supervised Multi-Frame Monocular DepthJamie Watson, Oisin Mac Aodha, Victor Prisacariu, Gabriel J. Brostow et al.CVPR 2021
- DepthInSpace: Exploitation and Fusion of Multiple Video Frames for Structured-Light Depth EstimationMohammad Mahdi Johari, Camilla Carta, François FleuretICCV 2021 · 12 citations
- Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth EstimationZhengming Zhou, Qiulei DongICCV 2023 · 16 citations
- Boosting Monocular Depth Estimation with Lightweight 3D Point FusionLam Huynh, Phong Nguyen, Jirí Matas, Esa Rahtu et al.ICCV 2021 · 32 citations
