SCAPO: Self-Supervised Category-Level Articulated Pose Estimation from a Single 3D Observation
Can Zhang, Gim Hee Lee
Abstract
Existing methods for category-level object articulation from a single 3D observation often rely on dense supervision, multi-frame inputs, or CAD templates, and still struggle to disentangle geometry from articulation or to recover explicit joint parameters. We propose SCAPO, a self-supervised framework that estimates canonical geometry, rigid part segmentation, and joint pivots, axes, and articulation states from a single RGB-D observation without ground-truth labels or category-specific models. Our SCAPO first uses an SE(3)-equivariant vector-neuron autoencoder to factor out global pose and align diverse instances into a shared canonical space. On this aligned shape, a joint-aware blend-skinning module is then designed to model part motion. We learn this representation through cycle reconstruction between observed and canonical shapes and cross-space alignment with a learnable canonical template that decouples shared category geometry from instance-specific residual shape. Experiments on synthetic and real articulated-object datasets show that our SCAPO recovers consistent part structure and accurate articulation parameters and outperforms all self-supervised baselines. Our source code is available at: https:// lulusindazc.github.io/SCAPOproject/.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5f18445a-76d5-4ab0-b2e0-aedd8513ec5eBuilds on19
- Vector Neurons: A General Framework for SO(3)-Equivariant NetworksCongyue Deng, Or Litany, Yueqi Duan, Adrien Poulenard et al.ICCV 2021 · 411 citations
- Neural Articulated Radiance FieldAtsuhiro Noguchi, Xiao Sun, Stephen Lin, Tatsuya HaradaICCV 2021 · 242 citations
- A-SDF: Learning Disentangled Signed Distance Functions for Articulated Shape RepresentationJiteng Mu, Weichao Qiu, Adam Kortylewski, Alan L. Yuille et al.ICCV 2021 · 138 citations
- HOI4D: A 4D Egocentric Dataset for Category-Level Human-Object InteractionYunze Liu, Yun Liu, Che Jiang, Kangbo Lyu et al.CVPR 2022 · 126 citations
- CAPTRA: CAtegory-level Pose Tracking for Rigid and Articulated Objects from Point CloudsYijia Weng, He Wang, Qiang Zhou, Yuzhe Qin et al.ICCV 2021 · 119 citations
Related papers
- Category-Level Articulated Object Pose EstimationXiaolong Li, He Wang, Li Yi, Leonidas J. Guibas et al.CVPR 2020
- Self-Supervised Category-Level Articulated Object Pose Estimation with Part-Level SE(3) EquivarianceXueyi Liu, Ji Zhang, Ruizhen Hu, Haibin Huang et al.ICLR 2023 · 3 citations
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas et al.NeurIPS 2021 · 61 citations
- Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the WildKaifeng Zhang, Yang Fu, Shubhankar Borse, Hong Cai et al.ICLR 2023 · 8 citations
- Unsupervised Volumetric AnimationAliaksandr Siarohin, Willi Menapace, Ivan Skorokhodov, Kyle Olszewski et al.CVPR 2023
