3D Equivariant Pose Regression via Direct Wigner-D Harmonics Prediction
Jongmin Lee, Minsu Cho
摘要
Determining the 3D orientations of an object in an image, known as single-image pose estimation, is a crucial task in 3D vision applications. Existing methods typically learn 3D rotations parametrized in the spatial domain using Euler angles or quaternions, but these representations often introduce discontinuities and singularities. SO(3)-equivariant networks enable the structured capture of pose patterns with data-efficient learning, but the parametrizations in spatial domain are incompatible with their architecture, particularly spherical CNNs, which operate in the frequency domain to enhance computational efficiency. To overcome these issues, we propose a frequency-domain approach that directly predicts Wigner-D coefficients for 3D rotation regression, aligning with the operations of spherical CNNs. Our SO(3)-equivariant pose harmonics predictor overcomes the limitations of spatial parameterizations, ensuring consistent pose estimation under arbitrary rotations. Trained with a frequency-domain regression loss, our method achieves state-of-the-art results on benchmarks such as ModelNet10-SO(3) and PASCAL3D+, with significant improvements in accuracy, robustness, and data efficiency.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Axis-Level Symmetry Detection with Group-Equivariant RepresentationWongyun Yu, Ahyun Seo, Minsu ChoICCV 2025 · 被引用 2 次
- Bridging Equivariant GNNs and Spherical CNNs for Structured Physical DomainsColin Kohler, Purvik Patel, Nathan Vaska, Justin A. Goodwin 等NeurIPS 2025 · 被引用 1 次
- RAVEN: End-to-end Equivariant Robot Learning with RGB CamerasDavid Klee, Boce Hu, Andrew Cole, Heng Tian 等ICLR 2026
- PSMix: Robust Point Cloud Recognition through Spectral Domain MixingXin Wei, Qin Yang, Hongji Zhao, Fei Gao 等ICML 2026
- REViT: Roto-reflection Equivariant Convolutional Vision TransformerSheir A. Zaheer, Alexander Holston, Chan Youn ParkICML 2026
它引用的顶会 Paper27
- CroCo: Self-Supervised Pre-training for 3D Vision Tasks by Cross-View CompletionPhilippe Weinzaepfel, Vincent Leroy, Thomas Lucas, Romain Brégier 等NeurIPS 2022 · 被引用 189 次
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 被引用 158 次
- An Analysis of SVD for Deep Rotation EstimationJake Levinson, Carlos Esteves, Kefan Chen, Noah Snavely 等NeurIPS 2020 · 被引用 131 次
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang 等ICLR 2024 · 被引用 126 次
- Reconstructing continuous distributions of 3D protein structure from cryo-EM imagesEllen D. Zhong, Tristan Bepler, Joseph H. Davis, Bonnie BergerICLR 2020 · 被引用 124 次
相关 Paper
- Image to Sphere: Learning Equivariant Features for Efficient Pose PredictionDavid Klee, Ondrej Biza, Robert Platt, Robin WaltersICLR 2023 · 被引用 3 次
- Equivariant Single View Pose Prediction Via Induced and Restriction RepresentationsOwen Howell, David Klee, Ondrej Biza, Linfeng Zhao 等NeurIPS 2023 · 被引用 4 次
- Learning to Orient Surfaces by Self-supervised Spherical CNNsRiccardo Spezialetti, Federico Stella, Marlon Marcon, Luciano Silva 等NeurIPS 2020 · 被引用 48 次
- Spin-Weighted Spherical CNNsCarlos Esteves, Ameesh Makadia, Kostas DaniilidisNeurIPS 2020 · 被引用 81 次
- VI-Net: Boosting Category-level 6D Object Pose Estimation via Learning Decoupled Rotations on the Spherical RepresentationsJiehong Lin, Zewei Wei, Yabin Zhang, Kui JiaICCV 2023 · 被引用 57 次
