Improving 360 Monocular Depth Estimation via Non-local Dense Prediction Transformer and Joint Supervised and Self-Supervised Learning
Ilwi Yun, Hyuk-Jae Lee, Chae-Eun Rhee
Abstract
Due to difficulties in acquiring ground truth depth of equirectangular (360 • ) images, the quality and quantity of equirectangular depth data today is insufficient to represent the various scenes in the world. Therefore, 360 • depth estimation studies, which relied solely on supervised learning, are destined to produce unsatisfactory results. Although self-supervised learning methods focusing on equirectangular images (EIs) are introduced, they often have incorrect or non-unique solutions, causing unstable performance. In this paper, we propose 360 • monocular depth estimation methods which improve on the areas that limited previous studies. First, we introduce a self-supervised 360 • depth learning method that only utilizes gravity-aligned videos, which has the potential to eliminate the needs for depth data during the training procedure. Second, we propose a joint learning scheme realized by combining supervised and self-supervised learning. The weakness of each learning is compensated, thus leading to more accurate depth estimation. Third, we propose a nonlocal fusion block, which can further retain the global information encoded by vision transformer when reconstructing the depths. With the proposed methods, we successfully apply the transformer to 360 • depth estimations, to the best of our knowledge, which has not been tried before. On several benchmarks, our approach achieves significant improvements over previous works and establishes a state of the art.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c0cdff05-159b-41b1-a9c0-130395855346Cited by top-tier papers6
- EGformer: Equirectangular Geometry-biased Transformer for 360 Depth EstimationIlwi Yun, Chanyong Shin, Hyunku Lee, Hyuk-Jae Lee et al.ICCV 2023 · 50 citations
- Pano-NeRF: Synthesizing High Dynamic Range Novel Views with Geometry from Sparse Low Dynamic Range Panoramic ImagesZhan Lu, Qian Zheng, Boxin Shi, Xudong JiangAAAI 2024 · 9 citations
- PanoPose: Self-supervised Relative Pose Estimation for Panoramic ImagesDiantao Tu, Hainan Cui, Xianwei Zheng, Shuhan ShenCVPR 2024 · 4 citations
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph OptimizationDongki Jung, Jaehoon Choi, Yonghan Lee, Dinesh ManochaNeurIPS 2025 · 4 citations
- EDM: Equirectangular Projection-Oriented Dense Kernelized Feature MatchingDongki Jung, Jaehoon Choi, Yonghan Lee, Somi Jeong et al.CVPR 2025
Builds on11
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Digging Into Self-Supervised Monocular Depth EstimationClément Godard, Oisin Mac Aodha, Michael Firman, Gabriel J. BrostowICCV 2019 · 2,416 citations
- Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown CamerasAriel Gordon, Hanhan Li, Rico Jonschkowski, Anelia AngelovaICCV 2019 · 397 citations
Related papers
- SliceNet: Deep Dense Depth Estimation From a Single Indoor Panorama Using a Slice-Based RepresentationGiovanni Pintore, Marco Agus, Eva Almansa, Jens Schneider et al.CVPR 2021
- OmniFusion: 360 Monocular Depth Estimation via Geometry-Aware FusionYuyan Li, Yuliang Guo, Zhixin Yan, Xinyu Huang et al.CVPR 2022 · 79 citations
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
- Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationNing-Hsu Wang, Yu-Lun LiuNeurIPS 2024 · 56 citations
- ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationOctave Mariotti, Oisin Mac Aodha, Hakan BilenICCV 2021 · 8 citations
