Image to Sphere: Learning Equivariant Features for Efficient Pose Prediction
David Klee, Ondrej Biza, Robert Platt, Robin Walters
摘要
Predicting the pose of objects from a single image is an important but difficult computer vision problem. Methods that predict a single point estimate do not predict the pose of objects with symmetries well and cannot represent uncertainty. Alternatively, some works predict a distribution over orientations in . However, training such models can be computation- and sample-inefficient. Instead, we propose a novel mapping of features from the image domain to the 3D rotation manifold. Our method then leverages equivariant layers, which are more sample efficient, and outputs a distribution over rotations that can be sampled at arbitrary resolution. We demonstrate the effectiveness of our method at object orientation prediction, and achieve state-of-the-art performance on the popular PASCAL3D+ dataset. Moreover, we show that our method can model complex object symmetries, without any modifications to the parameters or loss function. Code is available at https://dmklee.github.io/image2sphere.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Fourier Transporter: Bi-Equivariant Robotic Manipulation in 3DHaojie Huang, Owen Howell, Dian Wang, Xupeng Zhu 等ICLR 2024 · 被引用 40 次
- A General Theory of Correct, Incorrect, and Extrinsic EquivarianceDian Wang, Xupeng Zhu, Jung Yeon Park, Mingxi Jia 等NeurIPS 2023 · 被引用 23 次
- 3D Equivariant Visuomotor Policy Learning via Spherical ProjectionBoce Hu, Dian Wang, David Klee, Heng Tian 等NeurIPS 2025 · 被引用 9 次
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters 等ICLR 2026 · 被引用 9 次
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 被引用 6 次
它引用的顶会 Paper11
- SE(3)-Transformers: 3D Roto-Translation Equivariant Attention NetworksFabian Fuchs, Daniel E. Worrall, Volker Fischer, Max WellingNeurIPS 2020 · 被引用 1,025 次
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 被引用 486 次
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 被引用 482 次
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard 等ICCV 2021 · 被引用 411 次
- Equivariant Multi-View NetworksCarlos Esteves, Yinshuang Xu, Christine Allen-Blanchette, Kostas DaniilidisICCV 2019 · 被引用 108 次
相关 Paper
- Implicit-PDF: Non-Parametric Representation of Probability Distributions on the Rotation ManifoldKieran A. Murphy, Carlos Esteves, Varun Jampani, Srikumar Ramalingam 等ICML 2021 · 被引用 93 次
- Equivariant Single View Pose Prediction Via Induced and Restriction RepresentationsOwen Howell, David Klee, Ondrej Biza, Linfeng Zhao 等NeurIPS 2023 · 被引用 4 次
- E2PN: Efficient SE(3)-Equivariant Point NetworkMinghan Zhu, Maani Ghaffari, William A. Clark, Huei PengCVPR 2023
- Equivariant Descriptor Fields: SE(3)-Equivariant Energy-Based Models for End-to-End Visual Robotic Manipulation LearningHyunwoo Ryu, Hong-in Lee, Jeong-Hoon Lee, Jongeun ChoiICLR 2023 · 被引用 11 次
- Probabilistic Orientation Estimation with Matrix Fisher DistributionsDavid Mohlin, Josephine Sullivan, Gérald BianchiNeurIPS 2020 · 被引用 62 次
