Equivariant Ray Embeddings for Implicit Multi-View Depth Estimation
Yinshuang Xu, Dian Chen, Katherine Liu, Sergey Zakharov, Rares Ambrus, Kostas Daniilidis, Vitor Guizilini
Abstract
Incorporating inductive bias by embedding geometric entities (such as rays) as input has proven successful in multi-view learning. However, the methods adopting this technique typically lack equivariance, which is crucial for effective 3D learning. Equivariance serves as a valuable inductive prior, aiding in the generation of robust multi-view features for 3D scene understanding. In this paper, we explore the application of equivariant multi-view learning to depth estimation, not only recognizing its significance for computer vision and robotics but also addressing the limitations of previous research. Most prior studies have either overlooked equivariance in this setting or achieved only approximate equivariance through data augmentation, which often leads to inconsistencies across different reference frames. To address this issue, we propose to embed equivariance into the Perceiver IO architecture. We employ Spherical Harmonics for positional encoding to ensure 3D rotation equivariance, and develop a specialized equivariant encoder and decoder within the Perceiver IO architecture. To validate our model, we applied it to the task of stereo depth estimation, achieving state of the art results on real-world datasets without explicit geometric constraints or extensive data augmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Learning Generalizable Shape Completion with SIM(3) EquivarianceYuqing Wang, Zhaiyu Chen, Xiaoxiang ZhuNeurIPS 2025 · 2 citations
- Axis-Level Symmetry Detection with Group-Equivariant RepresentationWongyun Yu, Ahyun Seo, Minsu ChoICCV 2025 · 2 citations
Builds on31
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Reformer: The Efficient TransformerNikita Kitaev, Lukasz Kaiser, Anselm LevskayaICLR 2020 · 2,878 citations
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 1,432 citations
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- Depth Anything: Unleashing the Power of Large-Scale Unlabeled DataLihe Yang, Bingyi Kang, Zilong Huang, Xiaogang Xu et al.CVPR 2024 · 847 citations
Related papers
- Input-level Inductive Biases for 3D ReconstructionWang Yifan, Carl Doersch, Relja Arandjelovic, João Carreira et al.CVPR 2022 · 23 citations
- SE(3) Equivariant Convolution and Transformer in Ray SpaceYinshuang Xu, Jiahui Lei, Kostas DaniilidisNeurIPS 2023 · 6 citations
- Viewpoint Equivariance for Multi-View 3D Object DetectionDian Chen, Jie Li, Vitor Guizilini, Rares Ambrus et al.CVPR 2023
- Multi-body SE(3) Equivariance for Unsupervised Rigid Segmentation and Motion EstimationJia-Xing Zhong, Ta Ying Cheng, Yuhang He, Kai Lu et al.NeurIPS 2023 · 9 citations
- SO(3)-Equivariant ViT-Adapter for Data-Efficient Zero-Shot Sim-to-Real Indoor Panoramic Depth EstimationZiyan He, Qiudan Zhang, Lin Ma, Xu WangCVPR 2026
