Kick Back & Relax: Learning to Reconstruct the World by Watching SlowTV
Jaime Spencer, Simon Hadfield, Chris Russell, Richard Bowden
摘要
Self-supervised monocular depth estimation (SS-MDE) has the potential to scale to vast quantities of data. Unfortunately, existing approaches limit themselves to the automotive domain, resulting in models incapable of generalizing to complex environments such as natural or indoor settings.To address this, we propose a large-scale SlowTV dataset curated from YouTube, containing an order of magnitude more data than existing automotive datasets. SlowTV contains 1.7M images from a rich diversity of environments, such as worldwide seasonal hiking, scenic driving and scuba diving. Using this dataset, we train an SS-MDE model that provides zero-shot generalization to a large collection of indoor/outdoor datasets. The resulting model outperforms all existing SSL approaches and closes the gap on supervised SoTA, despite using a more efficient architecture.We additionally introduce a collection of best-practices to further maximize performance and zero-shot generalization. This includes 1) aspect ratio augmentation, 2) camera intrinsic estimation, 3) support frame randomization and 4) flexible motion estimation. Code is available at https://github.com/jspenmar/slowtv_monodepth.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- DUSt3R: Geometric 3D Vision Made EasyShuzhe Wang, Vincent Leroy, Yohann Cabon, Boris Chidlovskii 等CVPR 2024 · 被引用 302 次
- A Simple yet Universal Framework for Depth CompletionJin-Hwi Park, Hae-Gon JeonNeurIPS 2024 · 被引用 17 次
- Depth Pro: Sharp Monocular Metric Depth in Less Than a SecondAlexey Bochkovskiy, Amaël Delaunoy, Hugo Germain, Marcel Santos 等ICLR 2025 · 被引用 15 次
- Unsupervised Semantic Segmentation Through Depth-Guided Feature Correlation and SamplingLeon Sick, Dominik Engel, Pedro Hermosilla, Timo RopinskiCVPR 2024 · 被引用 6 次
- Depth Prompting for Sensor-Agnostic Depth EstimationJin-Hwi Park, Chanhwi Jeong, Junoh Lee, Hae-Gon JeonCVPR 2024 · 被引用 4 次
它引用的顶会 Paper24
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer 等CVPR 2022 · 被引用 6,782 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 被引用 2,647 次
相关 Paper
- Self-Supervised Human Depth Estimation From Monocular VideosFeitong Tan, Hao Zhu, Zhaopeng Cui, Siyu Zhu 等CVPR 2020
- MonoIndoor: Towards Good Practice of Self-Supervised Monocular Depth Estimation for Indoor EnvironmentsPan Ji, Runze Li, Bir Bhanu, Yi XuICCV 2021 · 被引用 82 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- MVSAnywhere: Zero-Shot Multi-View StereoSergio Izquierdo, Mohamed Sayed, Michael Firman, Guillermo Garcia-Hernando 等CVPR 2025
- Self-Supervised Monocular Trained Depth Estimation Using Self-Attention and Discrete Disparity VolumeAdrian Johnston, Gustavo CarneiroCVPR 2020
