3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction
Jongmin Lee, Minsu Cho
Abstract
Determining the 3D orientations of an object in an image, known as single-image pose estimation, is a crucial task in 3D vision applications. Existing methods typically learn 3D rotations parametrized in the spatial domain using Euler angles or quaternions, but these representations often introduce discontinuities and singularities. SO(3)-equivariant networks enable the structured capture of pose patterns with data-efficient learning, but the parametrizations in spatial domain are incompatible with their architecture, particularly spherical CNNs, which operate in the frequency domain to enhance computational efficiency. To overcome these issues, we propose a frequency-domain approach that directly predicts Wigner-D coefficients for 3D rotation regression, aligning with the operations of spherical CNNs. Our SO(3)-equivariant pose harmonics predictor overcomes the limitations of spatial parameterizations, ensuring consistent pose estimation under arbitrary rotations. Trained with a frequency-domain regression loss, our method achieves state-of-the-art results on benchmarks such as ModelNet10-SO(3) and PASCAL3D+, with significant improvements in accuracy, robustness, and data efficiency.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1d26e8be-4352-496c-9c15-ff9414b74a71Cited by top-tier papers6
- Axis-Level Symmetry Detection with Group-Equivariant RepresentationWongyun Yu, Ahyun Seo, Minsu ChoICCV 2025 · 2 citations
- Bridging Equivariant GNNs and Spherical CNNs for Structured Physical DomainsColin Kohler, Purvik Patel, Nathan Vaska, Justin A. Goodwin et al.NeurIPS 2025 · 1 citation
- RAVEN: End-to-end Equivariant Robot Learning with RGB CamerasDavid Klee, Boce Hu, Andrew Cole, Heng Tian et al.ICLR 2026
- PSMix: Robust Point Cloud Recognition through Spectral Domain MixingXin Wei, Qin Yang, Hongji Zhao, Fei Gao et al.ICML 2026
- REViT: Roto-reflection Equivariant Convolutional Vision TransformerSheir A. Zaheer, Alexander Holston, Chan Youn ParkICML 2026
Builds on27
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier et al.NeurIPS 2022 · 189 citations
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 158 citations
- An Analysis of SVD for Deep Rotation EstimationJake Levinson, Carlos Esteves, Kefan Chen, Noah Snavely et al.NeurIPS 2020 · 131 citations
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang et al.ICLR 2024 · 126 citations
- Reconstructing continuous distributions of 3D protein structure from cryo-EM imagesEllen D. Zhong, Tristan Bepler, Joseph H. Davis, Bonnie BergerICLR 2020 · 124 citations
Related papers
- Image to Sphere: Learning Equivariant Features for Efficient Pose PredictionDavid Klee, Ondrej Biza, Robert Platt, Robin WaltersICLR 2023 · 3 citations
- Equivariant Single View Pose Prediction Via Induced and Restriction RepresentationsOwen Howell, David Klee, Ondrej Biza, Linfeng Zhao et al.NeurIPS 2023 · 4 citations
- Learning to Orient Surfaces by Self-supervised Spherical CNNsRiccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva et al.NeurIPS 2020 · 48 citations
- Spin-Weighted Spherical CNNsCarlos Esteves, Ameesh Makadia, Kostas DaniilidisNeurIPS 2020 · 81 citations
- VI-Net: Boosting Category-level 6D Object Pose Estimation via Learning Decoupled Rotations on the Spherical RepresentationsJiehong Lin, Zewei Wei, Yabin Zhang, Kui JiaICCV 2023 · 57 citations
