ST360D: Spatial-to-Temporal Consistency for Training-free 360 Monocular Depth Estimation
Zidong Cao, Jinjing Zhu, Hao Ai, Lutao Jiang, Yuanhuiyi Lyu, Hui Xiong
摘要
• monocular depth estimation plays a crucial role in scene understanding owing to its 180 • × 360 • field-of-view (FoV). To mitigate the distortions brought by equirectangular projection, existing methods typically divide 360 • images into distortion-less perspective patches. However, since these patches are processed independently, depth inconsistencies are often introduced due to scale drift among patches. Recently, video depth estimation (VDE) models have leveraged temporal consistency for stable depth predictions across frames. Inspired by this, we propose to represent a 360 • image as a sequence of perspective frames, mimicking the viewpoint adjustments users make when exploring a 360 • scenario in virtual reality. Thus, the spatial consistency among perspective depth patches can be enhanced by exploiting the temporal consistency inherent in VDE models. To this end, we introduce a training-free pipeline for 360 • monocular depth estimation, called ST 2 360D. Specifically, ST 2 360D transforms a 360 • image into perspective video frames, predicts video depth maps using VDE models, and seamlessly merges these predictions into a complete 360 • depth map. To generate sequenced perspective frames that align with VDE models, we propose two tailored strategies. First, a spherical-uniform sampling (SUS) strategy is proposed to facilitate uniform sampling of perspective views across the sphere, avoiding oversampling in polar regions typically with limited structural details. Second, a latitude-guided scanning (LGS) strategy is introduced to organize the frames into a coherent sequence, starting from the equator, prioritizing low-latitude slices, and progressively moving toward higher latitudes. Extensive experiments demonstrate that ST 2 360D achieves strong zero-shot capability on several datasets, supporting resolutions up to 4K.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper32
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 被引用 6,759 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao 等NeurIPS 2024 · 被引用 2,305 次
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu 等CVPR 2024 · 被引用 847 次
- Consistent video depth estimationXuan Luo, Jia-Bin Huang, Richard Szeliski, Kevin Matzen 等SIGGRAPH 2020 · 被引用 321 次
相关 Paper
- Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationNing-Hsu Wang, Yu-Lun LiuNeurIPS 2024 · 被引用 56 次
- SVG: 3D Stereoscopic Video Generation via Denoising Frame MatrixPeng Dai, Feitong Tan, Qiangeng Xu, David Futschik 等ICLR 2025
- 360MonoDepth: High-Resolution 360° Monocular Depth EstimationManuel Rey-Area, Mingze Yuan, Christian RichardtCVPR 2022 · 被引用 80 次
- RPG360: Robust 360 Depth Estimation with Perspective Foundation Models and Graph OptimizationDongki Jung, Jaehoon Choi, Yonghan Lee, Dinesh ManochaNeurIPS 2025 · 被引用 4 次
- Beyond the Frame: Generating 360° Panoramic Videos from Perspective VideosRundong Luo, Matthew Wallingford, Ali Farhadi, Noah Snavely 等ICCV 2025 · 被引用 2 次
