SMOC-Net: Leveraging Camera Pose for Self-Supervised Monocular Object Pose Estimation
Tao Tan, Qiulei Dong
Abstract
Recently, self-supervised 6D object pose estimation, where synthetic images with object poses (sometimes jointly with un-annotated real images) are used for training, has attracted much attention in computer vision. Some typical works in literature employ a time-consuming differentiable renderer for object pose prediction at the training stage, so that (i) their performances on real images are generally limited due to the gap between their rendered images and real images and (ii) their training process is computationally expensive. To address the two problems, we propose a novel Network for Self-supervised Monocular Object pose estimation by utilizing the predicted Camera poses from un-annotated real images, called SMOC-Net. The proposed network is explored under a knowledge distillation framework, consisting of a teacher model and a student model. The teacher model contains a backbone estimation module for initial object pose estimation, and an object pose refiner for refining the initial object poses using a geometric constraint (called relative-pose constraint) derived from relative camera poses. The student model gains knowledge for object pose estimation from the teacher model by imposing the relative-pose constraint. Thanks to the relative-pose constraint, SMOC-Net could not only narrow the domain gap between synthetic and real data but also reduce the training cost. Experimental results on two public datasets demonstrate that SMOC-Net outperforms several state-of-the-art methods by a large margin while requiring much less training time than the differentiable-renderer-based methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 7513eedc-1a2b-410f-8b0e-be11e96f2b54Cited by top-tier papers2
- Environment-Agnostic Pose: Generating Environment-Independent Object Representations for 6D Pose EstimationShaobo Zhang, Yuhang Huang, Wanqing Zhao, Wei Zhao et al.ICCV 2025 · 3 citations
- ONDA-Pose: Occlusion-Aware Neural Domain Adaptation for Self-Supervised 6D Object Pose EstimationTao Tan, Qiulei DongCVPR 2025
Builds on14
- Soft Rasterizer: A Differentiable Renderer for Image-Based 3D ReasoningShichen Liu, Weikai Chen, Tianye Li, Hao LiICCV 2019 · 789 citations
- Pix2Pose: Pixel-Wise Coordinate Regression of Objects for 6D Pose EstimationKiru Park, Timothy Patten, Markus VinczeICCV 2019 · 527 citations
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- CDPN: Coordinates-Based Disentangled Pose Network for Real-Time RGB-Based 6-DoF Object Pose EstimationZhigang Li, Gu Wang, Xiangyang JiICCV 2019 · 482 citations
- EPro-PnP: Generalized End-to-End Probabilistic Perspective-n-Points for Monocular Object Pose EstimationHansheng Chen, Pichao Wang, Fan Wang, Wei Tian et al.CVPR 2022 · 175 citations
Related papers
- Pseudo Flow Consistency for Self-Supervised 6D Object Pose EstimationYang Hai, Rui Song, Jiaojiao Li, David Ferstl et al.ICCV 2023 · 13 citations
- DSC-PoseNet: Learning 6DoF Object Pose Estimation via Dual-Scale ConsistencyZongxin Yang, Xin Yu, Yi YangCVPR 2021
- Markerless Camera-to-Robot Pose Estimation via Self-Supervised Sim-to-Real TransferJingpei Lu, Florian Richter, Michael C. YipCVPR 2023
- Knowledge Distillation for 6D Pose Estimation by Aligning Distributions of Local PredictionsShuxuan Guo, Yinlin Hu, José M. Álvarez, Mathieu SalzmannCVPR 2023
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
