Learning Articulated Shape with Keypoint Pseudo-Labels from Web Images
Anastasis Stathopoulos, Georgios Pavlakos, Ligong Han, Dimitris N. Metaxas
Abstract
This paper shows that it is possible to learn models for monocular 3D reconstruction of articulated objects (e.g. horses, cows, sheep), using as few as 50-150 images labeled with 2D keypoints. Our proposed approach involves training category-specific keypoint estimators, generating 2D keypoint pseudo-labels on unlabeled web images, and using both the labeled and self-labeled sets to train 3D reconstruction models. It is based on two key insights: (1) 2D keypoint estimation networks trained on as few as 50-150 images of a given object category generalize well and generate reliable pseudo-labels; (2) a data selection mechanism can automatically create a "curated" subset of the unlabeled web images that can be used for training -we evaluate four data selection methods. Coupling these two insights enables us to train models that effectively utilize web images, resulting in improved 3D reconstruction performance for several articulated object categories beyond the fully-supervised baseline. Our approach can quickly bootstrap a model and requires only a few images labeled with 2D keypoints. This requirement can be easily satisfied for any new object category. To showcase the practicality of our approach for predicting the 3D shape of arbitrary object categories, we annotate 2D keypoints on 250 giraffe and bear images from COCO in just 2.5 hours per category.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a86e3aa-dc56-4915-a259-9c3bf3cba5b7Cited by top-tier papers2
- Score-Guided Diffusion for 3D Human RecoveryAnastasis Stathopoulos, Ligong Han, Dimitris N. MetaxasCVPR 2024 · 15 citations
- SAOR: Single-View Articulated Object ReconstructionMehmet Aygün, Oisin Mac AodhaCVPR 2024
Builds on15
- Learning to Reconstruct 3D Human Pose and Shape via Model-Fitting in the LoopNikos Kolotouros, Georgios Pavlakos, Michael J. Black, Kostas DaniilidisICCV 2019 · 1,139 citations
- PARE: Part Attention Regressor for 3D Human Body EstimationMuhammed Kocabas, Chun-Hao P. Huang, Otmar Hilliges, Michael J. BlackICCV 2021 · 509 citations
- Curriculum Labeling: Revisiting Pseudo-Labeling for Semi-Supervised LearningPaola Cascante-Bonilla, Fuwen Tan, Yanjun Qi, Vicente OrdonezAAAI 2021 · 362 citations
- Cross-Domain Adaptation for Animal Pose EstimationJinkun Cao, Hongyang Tang, Haoshu Fang, Xiaoyong Shen et al.ICCV 2019 · 209 citations
- Canonical Surface Mapping via Geometric Cycle ConsistencyNilesh Kulkarni, Shubham Tulsiani, Abhinav GuptaICCV 2019 · 104 citations
Related papers
- LASSIE: Learning Articulated Shapes from Sparse Image Ensemble via 3D Part DiscoveryChun-Han Yao, Wei-Chih Hung, Yuanzhen Li, Michael Rubinstein et al.NeurIPS 2022 · 83 citations
- LEPARD: Learning Explicit Part Discovery for 3D Articulated Shape ReconstructionDi Liu, Anastasis Stathopoulos, Qilong Zhangli, Yunhe Gao et al.NeurIPS 2023 · 24 citations
- DensePose 3D: Lifting Canonical Surface Maps of Articulated Objects to the Third DimensionRoman Shapovalov, David Novotný, Benjamin Graham, Patrick Labatut et al.ICCV 2021 · 10 citations
- ARTIC3D: Learning Robust Articulated 3D Shapes from Noisy Web Image CollectionsChun-Han Yao, Amit Raj, Wei-Chih Hung, Michael Rubinstein et al.NeurIPS 2023 · 25 citations
- C3DPO: Canonical 3D Pose Networks for Non-Rigid Structure From MotionDavid Novotný, Nikhila Ravi, Benjamin Graham, Natalia Neverova et al.ICCV 2019 · 126 citations
