Equivariant Single View Pose Prediction Via Induced and Restriction Representations
Owen Howell, David Klee, Ondrej Biza, Linfeng Zhao, Robin Walters
Abstract
Learning about the three-dimensional world from two-dimensional images is a fundamental problem in computer vision. An ideal neural network architecture for such tasks would leverage the fact that objects can be rotated and translated in three dimensions to make predictions about novel images. However, imposing SO(3)-equivariance on two-dimensional inputs is difficult because the group of three-dimensional rotations does not have a natural action on the two-dimensional plane. Specifically, it is possible that an element of SO(3) will rotate an image out of plane. We show that an algorithm that learns a three-dimensional representation of the world from two dimensional images must satisfy certain consistency properties which we formulate as SO(2)-equivariance constraints. We use the induced and restricted representations of SO(2) on SO(3) to construct and classify architectures which satisfy these consistency constraints. We prove that any architecture which respects said consistency constraints can be realized as an instance of our construction. We show that three previously proposed neural architectures for 3D pose prediction are special cases of our construction. We propose a new algorithm that is a learnable generalization of previously considered methods. We test our architecture on three pose predictions task and achieve SOTA results on both the PASCAL3D+ and SYMSOL pose estimation tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- 3D Equivariant Visuomotor Policy Learning via Spherical ProjectionBoce Hu, Dian Wang, David Klee, Heng Tian et al.NeurIPS 2025 · 9 citations
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters et al.ICLR 2026 · 9 citations
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 6 citations
- Cortical Policy: A Dual-Stream View Transformer for Robotic ManipulationXuening Zhang, Qi Lv, Xiang Deng, Miao Zhang et al.ICLR 2026 · 1 citation
- RAVEN: End-to-end Equivariant Robot Learning with RGB CamerasDavid Klee, Boce Hu, Andrew Cole, Heng Tian et al.ICLR 2026
Builds on17
- Reducing SO(3) Convolutions to SO(2) for Efficient Equivariant GNNsSaro Passaro, C. Lawrence ZitnickICML 2023 · 157 citations
- Gauge Equivariant Mesh CNNs: Anisotropic convolutions on geometric graphsPim de Haan, Maurice Weiler, Taco Cohen, Max WellingICLR 2021 · 139 citations
- Reconstructing continuous distributions of 3D protein structure from cryo-EM imagesEllen D. Zhong, Tristan Bepler, Joseph H. Davis, Bonnie BergerICLR 2020 · 124 citations
- Equivariant Multi-View NetworksCarlos Esteves, Yinshuang Xu, Christine Allen-Blanchette, Kostas DaniilidisICCV 2019 · 108 citations
- Implicit-PDF: Non-Parametric Representation of Probability Distributions on the Rotation ManifoldKieran A. Murphy, Carlos Esteves, Varun Jampani, Srikumar Ramalingam et al.ICML 2021 · 93 citations
Related papers
- Image to Sphere: Learning Equivariant Features for Efficient Pose PredictionDavid Klee, Ondrej Biza, Robert Platt, Robin WaltersICLR 2023 · 3 citations
- SE(3) Equivariant Convolution and Transformer in Ray SpaceYinshuang Xu, Jiahui Lei, Kostas DaniilidisNeurIPS 2023 · 6 citations
- Unsupervised Learning of Group Invariant and Equivariant RepresentationsRobin Winter, Marco Bertolini, Tuan Le, Frank Noé et al.NeurIPS 2022 · 61 citations
- Learning to Orient Surfaces by Self-supervised Spherical CNNsRiccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva et al.NeurIPS 2020 · 48 citations
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard et al.ICCV 2021 · 411 citations
