SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth Estimation
Ziyan He, Qiudan Zhang, Lin Ma, Xu Wang
Abstract
Panoramic depth estimation enables a complete understanding of 3D environments but faces significant challenges in generalizing to real-world scenes. While recent zero-shot depth models like Depth Anything achieve remarkable generalization on perspective images, their performance sharply degrades on panoramas due to projection distortions and the lack of spherical geometric awareness. Moreover, collecting large-scale panoramic RGB-D data is costly, hindering the large-scale training of panoramic foundation models. To address these issues, we propose an SO(3)-Equivariant ViT-Adapter, which transfers the powerful zero-shot capability of the perspective pre-trained ViT to panoramic depth estimation by explicitly incorporating a rotation-equivariant inductive bias. Our adapter introduces an SO(3) deformable cross-attention mechanism to effectively align SO(3)-equivariant features with perspective features, enhancing rotational consistency without modifying the ViT backbone. Trained solely on synthetic panoramas, our framework achieves robust zero-shot sim-to-real performance on real indoor benchmarks, including Matterport3D and Stanford2D3D, demonstrating both data efficiency and strong generalization for panoramic depth estimation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 67581703-e5c8-4158-93b8-af59a83d4a3aBuilds on35
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li et al.ICLR 2021 · 7,353 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
- AdaptFormer: Adapting Vision Transformers for Scalable Visual RecognitionShoufa Chen, Chongjian Ge, Zhan Tong, Jiangliu Wang et al.NeurIPS 2022 · 1,291 citations
- DAB-DETR: Dynamic Anchor Boxes are Better Queries for DETRShilong Liu, Feng Li, Hao Zhang, Xiao Yang et al.ICLR 2022 · 1,218 citations
Related papers
- DA2: Depth Anything in Any DirectionHaodong Li, Wangguandong Zheng, Jing He, Yuhao Liu et al.ICLR 2026 · 23 citations
- VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth EstimationJiayi Yuan, Haobo Jiang, De Wen Soh, Na ZhaoCVPR 2026 · 6 citations
- PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic ImageryYijing Guo, Mengjun Chao, Luo Wang, Tianyang Zhao et al.CVPR 2026 · 11 citations
- Bending Reality: Distortion-aware Transformers for Adapting to Panoramic Semantic SegmentationJiaming Zhang, Kailun Yang, Chaoxiang Ma, Simon Reiß et al.CVPR 2022 · 100 citations
- Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationNing-Hsu Wang, Yu-Lun LiuNeurIPS 2024 · 56 citations
