SfM-TTR: Using Structure from Motion for Test-Time Refinement of Single-View Depth Networks
Sergio Izquierdo, Javier Civera
摘要
Estimating a dense depth map from a single view is geometrically ill-posed, and state-of-the-art methods rely on learning depth's relation with visual appearance using deep neural networks. On the other hand, Structure from Motion (SfM) leverages multi-view constraints to produce very accurate but sparse maps, as matching across images is typically limited by locally discriminative texture. In this work, we combine the strengths of both approaches by proposing a novel test-time refinement (TTR) method, denoted as SfM-TTR, that boosts the performance of single-view depth networks at test time using SfM multi-view cues. Specifically, and differently from the state of the art, we use sparse SfM point clouds as test-time self-supervisory signal, fine-tuning the network encoder to learn a better representation of the test scene. Our results show how the addition of SfM-TTR to several state-of-the-art self-supervised and supervised networks improves significantly their performance, outperforming previous TTR baselines mainly based on photometric multi-view consistency. The code is available at https://github.com/serizba/SfM-TTR .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- LightDepth: Single-View Depth Self-Supervision from Illumination DeclineJavier Rodriguez Puigvert, Victor M. Batlle, J. M. M. Montiel, Ruben Martinez-Cantin 等ICCV 2023 · 被引用 22 次
- MVSAnywhere: Zero-Shot Multi-View StereoSergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando 等CVPR 2025
- Layered Motion Fusion: Lifting Motion Segmentation to 3D in Egocentric VideosVadim Tschernezki, Diane Larlus, Iro Laina, Andrea VedaldiCVPR 2025
它引用的顶会 Paper12
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 被引用 2,416 次
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 被引用 397 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
- Neural Window Fully-connected CRFs for Monocular Depth EstimationWeihao Yuan, Xiaodong Gu, Zuozhuo Dai, Siyu Zhu 等CVPR 2022 · 被引用 320 次
- Self-Supervised Learning With Geometric Constraints in Monocular Video: Connecting Flow, Depth, and CameraYuhua Chen, Cordelia Schmid, Cristian SminchisescuICCV 2019 · 被引用 265 次
相关 Paper
- Test3R: Learning to Reconstruct 3D at Test TimeYuheng Yuan, Qiuhong Shen, Shizun Wang, Xingyi Yang 等NeurIPS 2025 · 被引用 19 次
- The Temporal Opportunist: Self-Supervised Multi-Frame Monocular DepthJamie Watson, Oisin Mac Aodha, Victor Prisacariu, Gabriel J. Brostow 等CVPR 2021
- DepthInSpace: Exploitation and Fusion of Multiple Video Frames for Structured-Light Depth EstimationMohammad Mahdi Johari, Camilla Carta, François FleuretICCV 2021 · 被引用 12 次
- Two-in-One Depth: Bridging the Gap Between Monocular and Binocular Self-supervised Depth EstimationZhengming Zhou, Qiulei DongICCV 2023 · 被引用 16 次
- Boosting Monocular Depth Estimation with Lightweight 3D Point FusionLam Huynh, Phong Nguyen, Jirí Matas, Esa Rahtu 等ICCV 2021 · 被引用 32 次
