SE(3) Equivariant Convolution and Transformer in Ray Space
Yinshuang Xu, Jiahui Lei, Kostas Daniilidis
Abstract
3D reconstruction and novel view rendering can greatly benefit from geometric priors when the input views are not sufficient in terms of coverage and inter-view baselines. Deep learning of geometric priors from 2D images requires each image to be represented in a 2D canonical frame and the prior to be learned in a given or learned 3D canonical frame. In this paper, given only the relative poses of the cameras, we show how to learn priors from multiple views equivariant to coordinate frame transformations by proposing an SE(3)-equivariant convolution and transformer in the space of rays in 3D. We model the ray space as a homogeneous space of SE(3) and introduce the SE(3)-equivariant convolution in ray space. Depending on the output domain of the convolution, we present convolution-based SE(3)-equivariant maps from ray space to ray space and to R 3 . Our mathematical framework allows us to go beyond convolution to SE(3)-equivariant attention in the ray space. We showcase how to tailor and adapt the equivariant convolution and transformer in the tasks of equivariant 3D reconstruction and equivariant neural rendering from multiple views. We demonstrate SE(3)-equivariance by obtaining robust results in roto-translated datasets without performing transformation augmentation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Equivariant Ray Embeddings for Implicit Multi-View Depth EstimationYinshuang Xu, Dian Chen, Katherine Liu, Sergey Zakharov et al.NeurIPS 2024 · 11 citations
- EqNIO: Subequivariant Neural Inertial OdometryRoyina Karegoudra Jayanth, Yinshuang Xu, Ziyun Wang, Evangelos Chatzipantazis et al.ICLR 2025
Builds on34
- E(n) Equivariant Graph Neural NetworksVictor Garcia Satorras, Emiel Hoogeboom, Max WellingICML 2021 · 1,432 citations
- PIFu: Pixel-Aligned Implicit Function for High-Resolution Clothed Human DigitizationShunsuke Saito, Zeng Huang, Ryota Natsume, Shigeo Morishima et al.ICCV 2019 · 1,411 citations
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- MVSNeRF: Fast Generalizable Radiance Field Reconstruction from Multi-View StereoAnpei Chen, Zexiang Xu, Fuqiang Zhao, Xiaoshuai Zhang et al.ICCV 2021 · 1,024 citations
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum et al.NeurIPS 2021 · 426 citations
Related papers
- Pose-Transformed Equivariant Network for 3D Point Trajectory PredictionRuixuan Yu, Jian SunCVPR 2024 · 2 citations
- SE(3)-bi-equivariant Transformers for Point Cloud AssemblyZiming Wang, Rebecka JörnstenNeurIPS 2024 · 5 citations
- Gaussian Process Priors for View-Aware InferenceYuxin Hou, Ari Heljakka, Arno SolinAAAI 2021 · 1 citation
- Equivariant Single View Pose Prediction Via Induced and Restriction RepresentationsOwen Howell, David Klee, Ondrej Biza, Linfeng Zhao et al.NeurIPS 2023 · 4 citations
- Learning Coordinate-based Convolutional Kernels for Continuous SE(3) Equivariant and Efficient Point Cloud AnalysisJaein Kim, Hee Bin Yoo, Dong-Sig Han, Byoung-Tak ZhangCVPR 2026
