Elite360D: Towards Efficient 360 Depth Estimation via Semantic- and Distance-Aware Bi-Projection Fusion
Hao Ai, Lin Wang
摘要
360 depth estimation has recently received great attention for 3D reconstruction owing to its omnidirectional field of view (FoV). Recent approaches are predominantly focused on cross-projection fusion with geometry-based reprojection: they fuse 360 images with equirectangular projection (ERP) and another projection type, e.g., cubemap projection to estimate depth with the ERP format. However, these methods suffer from 1) limited local receptive fields, making it hardly possible to capture large FoV scenes, and 2) prohibitive computational cost, caused by the complex cross-projection fusion module design. In this paper, we propose Elite360D, a novel framework that inputs the ERP image and icosahedron projection (ICOSAP) point set, which is undistorted and spatially continuous. Elite360D is superior in its capacity in learning a representation from a local-with-global perspective. With a flexible ERP image encoder, it includes an ICOSAP point encoder, and a Biprojection Bi-attention Fusion (B2F) module (totally 1M parameters). Specifically, the ERP image encoder can take various perspective image-trained backbones (e.g., ResNet, Transformer) to extract local features. The point encoder extracts the global features from the ICOSAP. Then, the B2F module captures the semantic- and distance-aware dependencies between each pixel of the ERP feature and the entire ICOSAP feature set. Without specific backbone design and obvious computational cost increase, Elite360D outperforms the prior arts on several benchmark datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- Depth Any Panoramas: A Foundation Model for Panoramic Depth EstimationXin Lin, Meixi Song, Dizhe Zhang, Wenxuan Lu 等CVPR 2026 · 被引用 27 次
- DA2: Depth Anything in Any DirectionHaodong Li, Wangguandong Zheng, Jing He, Yuhao Liu 等ICLR 2026 · 被引用 23 次
- PanoVGGT: Feed-Forward 3D Reconstruction from Panoramic ImageryYijing Guo, Mengjun Chao, Luo Wang, Tianyang Zhao 等CVPR 2026 · 被引用 11 次
- One2Scene: Geometric Consistent Explorable 3D Scene Generation from a Single ImagePengfei Wang, Liyi Chen, Zhiyuan Ma, Yanjun Guo 等ICLR 2026 · 被引用 11 次
- VGGT-360: Geometry-Consistent Zero-Shot Panoramic Depth EstimationJiayi Yuan, Haobo Jiang, De Wen Soh, Na ZhaoCVPR 2026 · 被引用 6 次
它引用的顶会 Paper21
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Rope3D: The Roadside Perception Dataset for Autonomous Driving and Monocular 3D Object Detection TaskXiaoqing Ye, Mao Shu, Hanyu Li, Yifeng Shi 等CVPR 2022 · 被引用 130 次
- Orientation-Aware Semantic Segmentation on Icosahedron SpheresChao Zhang, Stephan Liwicki, William Smith, Roberto CipollaICCV 2019 · 被引用 90 次
- 360MonoDepth: High-Resolution 360° Monocular Depth EstimationManuel Rey-Area, Mingze Yuan, Christian RichardtCVPR 2022 · 被引用 80 次
相关 Paper
- BiFuse: Monocular 360 Depth Estimation via Bi-Projection FusionFu-En Wang, Yu-Hsuan Yeh, Min Sun, Wei-Chen Chiu 等CVPR 2020
- SphereUFormer: A U-Shaped Transformer for Spherical 360 PerceptionYaniv Benny, Lior WolfCVPR 2025
- Depth Anywhere: Enhancing 360 Monocular Depth Estimation via Perspective Distillation and Unlabeled Data AugmentationNing-Hsu Wang, Yu-Lun LiuNeurIPS 2024 · 被引用 56 次
- SphereSR: 360° Image Super-Resolution with Arbitrary Projection via Continuous Spherical Image RepresentationYoungho Yoon, Inchul Chung, Lin Wang, Kuk-Jin YoonCVPR 2022 · 被引用 44 次
- Omnidirectional Image Super-resolution via Bi-projection FusionJiangang Wang, Yuning Cui, Yawen Li, Wenqi Ren 等AAAI 2024 · 被引用 15 次
