Image to Sphere: Learning Equivariant Features for Efficient Pose Prediction
David Klee, Ondrej Biza, Robert Platt, Robin Walters
Abstract
Predicting the pose of objects from a single image is an important but difficult computer vision problem. Methods that predict a single point estimate do not predict the pose of objects with symmetries well and cannot represent uncertainty. Alternatively, some works predict a distribution over orientations in . However, training such models can be computation- and sample-inefficient. Instead, we propose a novel mapping of features from the image domain to the 3D rotation manifold. Our method then leverages equivariant layers, which are more sample efficient, and outputs a distribution over rotations that can be sampled at arbitrary resolution. We demonstrate the effectiveness of our method at object orientation prediction, and achieve state-of-the-art performance on the popular PASCAL3D+ dataset. Moreover, we show that our method can model complex object symmetries, without any modifications to the parameters or loss function. Code is available at https://dmklee.github.io/image2sphere.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e7a6a0e2-03ca-4441-b6dc-bbfe685c73c2Cited by top-tier papers13
- Fourier Transporter: Bi-Equivariant Robotic Manipulation in 3DHaojie Huang, Owen Howell, Dian Wang, Xupeng Zhu et al.ICLR 2024 · 40 citations
- A General Theory of Correct, Incorrect, and Extrinsic EquivarianceDian Wang, Xupeng Zhu, Jung Yeon Park, Mingxi Jia et al.NeurIPS 2023 · 23 citations
- 3D Equivariant Visuomotor Policy Learning via Spherical ProjectionBoce Hu, Dian Wang, David Klee, Heng Tian et al.NeurIPS 2025 · 9 citations
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters et al.ICLR 2026 · 9 citations
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 6 citations
Builds on11
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 1,025 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard et al.ICCV 2021 · 411 citations
- Equivariant Multi-View NetworksCarlos Esteves, Yinshuang Xu, Christine Allen-Blanchette, Kostas DaniilidisICCV 2019 · 108 citations
Related papers
- Implicit-PDF: Non-Parametric Representation of Probability Distributions on the Rotation ManifoldKieran A. Murphy, Carlos Esteves, Varun Jampani, Srikumar Ramalingam et al.ICML 2021 · 93 citations
- Equivariant Single View Pose Prediction Via Induced and Restriction RepresentationsOwen Howell, David Klee, Ondrej Biza, Linfeng Zhao et al.NeurIPS 2023 · 4 citations
- E2PN: Efficient SE(3)-Equivariant Point NetworkMinghan Zhu, Maani Ghaffari, William A. Clark, Huei PengCVPR 2023
- Equivariant Descriptor Fields: SE(3)-Equivariant Energy-Based Models for End-to-End Visual Robotic Manipulation LearningHyunwoo Ryu, Hong-in Lee, Jeong-Hoon Lee, Jongeun ChoiICLR 2023 · 11 citations
- Probabilistic Orientation Estimation with Matrix Fisher DistributionsDavid Mohlin, Josephine Sullivan, Gérald BianchiNeurIPS 2020 · 62 citations
