VI-Net: Boosting Category-level 6D Object Pose Estimation via Learning Decoupled Rotations on the Spherical Representations
Jiehong Lin, Zewei Wei, Yabin Zhang, Kui Jia
Abstract
Rotation estimation of high precision from an RGB-D object observation is a huge challenge in 6D object pose estimation, due to the difficulty of learning in the non-linear space of SO (3). In this paper, we propose a novel rotation estimation network, termed as VI-Net, to make the task easier by decoupling the rotation as the combination of a viewpoint rotation and an in-plane rotation. More specifically, VI-Net bases the feature learning on the sphere with two individual branches for the estimates of two factorized rotations, where a V-Branch is employed to learn the viewpoint rotation via binary classification on the spherical signals, while another I-Branch is used to estimate the in-plane rotation by transforming the signals to view from the zenith direction. To process the spherical signals, a Spherical Feature Pyramid Network is constructed based on a novel design of SPAtial Spherical Convolution (SPA-SConv), which settles the boundary problem of spherical signals via feature padding and realizes viewpoint-equivariant feature extraction by symmetric convolutional operations. We apply the proposed VI-Net to the challenging task of category-level 6D object pose estimation for predicting the poses of unknown objects without available CAD models; experiments on the benchmarking datasets confirm the efficacy of our method, which outperforms the existing ones with a large margin in the regime of high precision.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext df0cc489-0ee9-4360-ba37-48c4530e2a25Cited by top-tier papers18
- Vision Foundation Model Enables Generalizable Object Pose EstimationKai Chen, Yiyao Ma, Xingyu Lin, Stephen James et al.NeurIPS 2024 · 5 citations
- RFMPose: Generative Category-level Object Pose Estimation via Riemannian Flow MatchingWenzhe Ouyang, Qi Ye, Jinghua Wang, Zenglin Xu et al.NeurIPS 2025 · 5 citations
- ComPose: A Unified Completion-Pose Framework for Robust Category-Level Object Pose EstimationHuan Ren, Yihan Chen, Chuxin Wang, Nailong Liu et al.CVPR 2026 · 4 citations
- CleanPose: Category-Level Object Pose Estimation via Causal Learning and Knowledge DistillationXiao Lin, Yun Peng, Liuyi Wang, Xianyou Zhong et al.ICCV 2025 · 3 citations
- KeyPose: Category-Level 6D Object Pose Estimation with Self-Adaptive KeypointsSheng Yu, Di-Hua Zhai, Yuanqing XiaAAAI 2025 · 2 citations
Builds on11
- 6-DOF GraspNet: Variational Grasp Generation for Object ManipulationArsalan Mousavian, Clemens Eppner, Dieter FoxICCV 2019 · 673 citations
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 183 citations
- DualPoseNet: Category-level 6D Object Pose and Size Estimation Using Dual Pose Network with Refined Learning of Pose ConsistencyJiehong Lin, Zewei Wei, Zhihao Li, Songcen Xu et al.ICCV 2021 · 169 citations
- GPV-Pose: Category-level Object Pose Estimation via Geometry-guided Point-wise VotingYan Di, Ruida Zhang, Zhiqiang Lou, Fabian Manhardt et al.CVPR 2022 · 141 citations
- VISTA: Boosting 3D Object Detection via Dual Cross-VIew SpaTial AttentionShengheng Deng, Zhihao Liang, Lin Sun, Kui JiaCVPR 2022 · 92 citations
Related papers
- Sparse Steerable Convolutions: An Efficient Learning of SE(3)-Equivariant Features for Estimation and Tracking of Object Poses in 3D SpaceJiehong Lin, Hongyang Li, Ke Chen, Jiangbo Lu et al.NeurIPS 2021 · 28 citations
- Learning Shape-Independent Transformation via Spherical Representations for Category-Level Object Pose EstimationHuan Ren, Wenfei Yang, Xiang Liu, Shifeng Zhang et al.ICLR 2025
- G2L-Net: Global to Local Network for Real-Time 6D Pose Estimation With Embedding Vector FeaturesWei Chen, Xi Jia, Hyung Jin Chang, Jinming Duan et al.CVPR 2020
- FS-Net: Fast Shape-Based Network for Category-Level 6D Object Pose Estimation With Decoupled Rotation MechanismWei Chen, Xi Jia, Hyung Jin Chang, Jinming Duan et al.CVPR 2021
- Cylindrical Convolutional Networks for Joint Object Detection and Viewpoint EstimationSunghun Joung, Seungryong Kim, Hanjae Kim, Minsu Kim et al.CVPR 2020
