ST360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth Estimation
Zidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang, Yuanhuiyi Lyu, Hui Xiong
Abstract
• monocular depth estimation plays a crucial role in scene understanding owing to its 180 • × 360 • field-of-view (FoV). To mitigate the distortions brought by equirectangular projection, existing methods typically divide 360 • images into distortion-less perspective patches. However, since these patches are processed independently, depth inconsistencies are often introduced due to scale drift among patches. Recently, video depth estimation (VDE) models have leveraged temporal consistency for stable depth predictions across frames. Inspired by this, we propose to represent a 360 • image as a sequence of perspective frames, mimicking the viewpoint adjustments users make when exploring a 360 • scenario in virtual reality. Thus, the spatial consistency among perspective depth patches can be enhanced by exploiting the temporal consistency inherent in VDE models. To this end, we introduce a training-free pipeline for 360 • monocular depth estimation, called ST 2 360D. Specifically, ST 2 360D transforms a 360 • image into perspective video frames, predicts video depth maps using VDE models, and seamlessly merges these predictions into a complete 360 • depth map. To generate sequenced perspective frames that align with VDE models, we propose two tailored strategies. First, a spherical-uniform sampling (SUS) strategy is proposed to facilitate uniform sampling of perspective views across the sphere, avoiding oversampling in polar regions typically with limited structural details. Second, a latitude-guided scanning (LGS) strategy is introduced to organize the frames into a coherent sequence, starting from the equator, prioritizing low-latitude slices, and progressively moving toward higher latitudes. Extensive experiments demonstrate that ST 2 360D achieves strong zero-shot capability on several datasets, supporting resolutions up to 4K.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext dc62bfb1-afde-4c2a-becb-3008a2d7123bBuilds on32
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen et al.SIGGRAPH 2020 · 321 citations
Related papers
- Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationNing-Hsu Wang, Yu-Lun LiuNeurIPS 2024 · 56 citations
- SVG: 3D Stereoscopic Video Generation via Denoising Frame MatrixPeng Dai, Feitong Tan, Qiangeng Xu, David Futschik et al.ICLR 2025
- 360MonoDepth: High-Resolution 360° Monocular Depth EstimationManuel Rey-Area, Mingze Yuan, Christian RichardtCVPR 2022 · 80 citations
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph OptimizationDongki Jung, Jaehoon Choi, Yonghan Lee, Dinesh ManochaNeurIPS 2025 · 4 citations
- Beyond the Frame: Generating 360° Panoramic Videos from Perspective VideosRundong Luo, Matthew Wallingford, Ali Farhadi, Noah Snavely et al.ICCV 2025 · 2 citations
