Equivariant Single View Pose Prediction Via Induced and Restriction Representations
Owen Howell, David Klee, Ondrej Biza, Linfeng Zhao, Robin Walters
摘要
Learning about the three-dimensional world from two-dimensional images is a fundamental problem in computer vision. An ideal neural network architecture for such tasks would leverage the fact that objects can be rotated and translated in three dimensions to make predictions about novel images. However, imposing SO(3)-equivariance on two-dimensional inputs is difficult because the group of three-dimensional rotations does not have a natural action on the two-dimensional plane. Specifically, it is possible that an element of SO(3) will rotate an image out of plane. We show that an algorithm that learns a three-dimensional representation of the world from two dimensional images must satisfy certain consistency properties which we formulate as SO(2)-equivariance constraints. We use the induced and restricted representations of SO(2) on SO(3) to construct and classify architectures which satisfy these consistency constraints. We prove that any architecture which respects said consistency constraints can be realized as an instance of our construction. We show that three previously proposed neural architectures for 3D pose prediction are special cases of our construction. We propose a new algorithm that is a learnable generalization of previously considered methods. We test our architecture on three pose predictions task and achieve SOTA results on both the PASCAL3D+ and SYMSOL pose estimation tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- 3D Equivariant Visuomotor Policy Learning via Spherical ProjectionBoce Hu, Dian Wang, David Klee, Heng Tian 等NeurIPS 2025 · 被引用 9 次
- EquAct: An SE(3)-Equivariant Multi-Task Transformer for 3D Robotic ManipulationXupeng Zhu, Yu Qi, Yizhe Zhu, Robin Walters 等ICLR 2026 · 被引用 9 次
- 3D Equivariant Pose Regression via Direct Wigner-D Harmonics PredictionJongmin Lee, Minsu ChoNeurIPS 2024 · 被引用 6 次
- Cortical Policy: A Dual-Stream View Transformer for Robotic ManipulationXuening Zhang, Qi Lv, Xiang Deng, Miao Zhang 等ICLR 2026 · 被引用 1 次
- RAVEN: End-to-end Equivariant Robot Learning with RGB CamerasDavid Klee, Boce Hu, Andrew Cole, Heng Tian 等ICLR 2026
它引用的顶会 Paper17
- Reducing SO(3) Convolutions to SO(2) for Efficient Equivariant GNNsSaro Passaro, C. Lawrence ZitnickICML 2023 · 被引用 157 次
- Gauge Equivariant Mesh CNNs: Anisotropic convolutions on geometric graphsPim de Haan, Maurice Weiler, Taco Cohen, Max WellingICLR 2021 · 被引用 139 次
- Reconstructing continuous distributions of 3D protein structure from cryo-EM imagesEllen D. Zhong, Tristan Bepler, Joseph H. Davis, Bonnie BergerICLR 2020 · 被引用 124 次
- Equivariant Multi-View NetworksCarlos Esteves, Yinshuang Xu, Christine Allen-Blanchette, Kostas DaniilidisICCV 2019 · 被引用 108 次
- Implicit-PDF: Non-Parametric Representation of Probability Distributions on the Rotation ManifoldKieran A. Murphy, Carlos Esteves, Varun Jampani, Srikumar Ramalingam 等ICML 2021 · 被引用 93 次
相关 Paper
- Image to Sphere: Learning Equivariant Features for Efficient Pose PredictionDavid Klee, Ondrej Biza, Robert Platt, Robin WaltersICLR 2023 · 被引用 3 次
- SE(3) Equivariant Convolution and Transformer in Ray SpaceYinshuang Xu, Jiahui Lei, Kostas DaniilidisNeurIPS 2023 · 被引用 6 次
- Unsupervised Learning of Group Invariant and Equivariant RepresentationsRobin Winter, Marco Bertolini, Tuan Le, Frank Noé 等NeurIPS 2022 · 被引用 61 次
- Learning to Orient Surfaces by Self-supervised Spherical CNNsRiccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva 等NeurIPS 2020 · 被引用 48 次
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard 等ICCV 2021 · 被引用 411 次
