Self-Supervised 3D Human Pose Estimation via Part Guided Novel Image Synthesis
Jogendra Nath Kundu, Siddharth Seth, Varun Jampani, Mugalodi Rakesh, R. Venkatesh Babu, Anirban Chakraborty
Abstract
Camera captured human pose is an outcome of several sources of variation. Performance of supervised 3D pose estimation approaches comes at the cost of dispensing with variations, such as shape and appearance, that may be useful for solving other related tasks. As a result, the learned model not only inculcates task-bias but also dataset-bias because of its strong reliance on the annotated samples, which also holds true for weakly-supervised models. Acknowledging this, we propose a self-supervised learning framework 1 to disentangle such variations from unlabeled video frames. We leverage the prior knowledge on human skeleton and poses in the form of a single part based 2D puppet model, human pose articulation constraints, and a set of unpaired 3D poses. Our differentiable formalization, bridging the representation gap between the 3D pose and spatial part maps, not only facilitates discovery of interpretable pose disentanglement, but also allows us to operate on videos with diverse camera movements. Qualitative results on unseen in-the-wild datasets establish our superior generalization across multiple tasks beyond the primary tasks of 3D pose estimation and part segmentation. Furthermore, we demonstrate state-of-the-art weaklysupervised 3D pose estimation performance on both Hu-man3.6M and MPI-INF-3DHP datasets.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers18
- Uncertainty-Aware Adaptation for Self-Supervised 3D Human Pose EstimationJogendra Nath Kundu, Siddharth Seth, Pradyumna YM, Varun Jampani et al.CVPR 2022 · 41 citations
- PoseTriplet: Co-evolving 3D Human Pose Estimation, Imitation, and Hallucination under Self-supervisionKehong Gong, Bingbing Li, Jianfeng Zhang, Tao Wang et al.CVPR 2022 · 40 citations
- Towards Alleviating the Modeling Ambiguity of Unsupervised Monocular 3D Human Pose EstimationZhenbo Yu, Bingbing Ni, Jingwei Xu, Junjie Wang et al.ICCV 2021 · 39 citations
- AdaptPose: Cross-Dataset Adaptation for 3D Human Pose Estimation by Learnable Motion GenerationMohsen Gholami, Bastian Wandt, Helge Rhodin, Rabab Ward et al.CVPR 2022 · 31 citations
- MetaPose: Fast 3D Pose from Multiple Views without 3D SupervisionBen Usman, Andrea Tagliasacchi, Kate Saenko, Avneesh SudCVPR 2022 · 29 citations
Builds on3
- C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From MotionDavid Novotný, Nikhila Ravi, Benjamin Graham, Natalia Neverova et al.ICCV 2019 · 126 citations
- Deep Non-Rigid Structure From MotionChen Kong, Simon LuceyICCV 2019 · 72 citations
- Kinematic-Structure-Preserved Representation for Unsupervised 3D Human Pose EstimationJogendra Nath Kundu, Siddharth Seth, Rahul M. V., Mugalodi Rakesh et al.AAAI 2020 · 57 citations
Related papers
- CanonPose: Self-Supervised Monocular 3D Human Pose Estimation in the WildBastian Wandt, Marco Rudolph, Petrissa Zell, Helge Rhodin et al.CVPR 2021
- Self-Supervised Learning of Interpretable Keypoints From Unlabelled VideosTomas Jakab, Ankush Gupta, Hakan Bilen, Andrea VedaldiCVPR 2020
- Category-Level Articulated Object Pose EstimationXiaolong Li, He Wang, Li Yi, Leonidas J. Guibas et al.CVPR 2020
- Self-Supervised Category-Level Articulated Object Pose Estimation with Part-Level SE(3) EquivarianceXueyi Liu, Ji Zhang, Ruizhen Hu, Haibin Huang et al.ICLR 2023 · 3 citations
- Video Autoencoder: self-supervised disentanglement of static 3D structure and motionZihang Lai, Sifei Liu, Alexei A. Efros, Xiaolong WangICCV 2021 · 37 citations
