Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D Pose
Angtian Wang, Shenxiao Mei, Alan L. Yuille, Adam Kortylewski
Abstract
We study the problem of learning to estimate the 3D object pose from a few labelled examples and a collection of unlabelled data. Our main contribution is a learning framework, neural view synthesis and matching, that can transfer the 3D pose annotation from the labelled to unlabelled images reliably, despite unseen 3D views and nuisance variations such as the object shape, texture, illumination or scene context. In our approach, objects are represented as 3D cuboid meshes composed of feature vectors at each mesh vertex. The model is initialized from a few labelled images and is subsequently used to synthesize feature representations of unseen 3D views. The synthesized views are matched with the feature representations of unlabelled images to generate pseudo-labels of the 3D pose. The pseudo-labelled data is, in turn, used to train the feature extractor such that the features at each mesh vertex are more invariant across varying 3D views of the object. Our model is trained in an EM-type manner alternating between increasing the 3D pose invariance of the feature extractor and annotating unlabelled data through neural view synthesis and matching. We demonstrate the effectiveness of the proposed semi-supervised learning framework for 3D pose estimation on the PASCAL3D+ and KITTI datasets. We find that our approach outperforms all baselines by a wide margin, particularly in an extreme few-shot setting where only 7 annotated images are given. Remarkably, we observe that our model also achieves an exceptional robustness in out-of-distribution scenarios that involve partial occlusion. The code is available at https://github.com/Angtian/NeuralVS .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c67a3ea8-d5ef-4e74-93f1-afff1c2d1832Cited by top-tier papers2
- Generating Images with 3D Annotations Using Diffusion ModelsWufei Ma, Qihao Liu, Jiahao Wang, Angtian Wang et al.ICLR 2024 · 18 citations
- FisherMatch: Semi-Supervised Rotation Regression via Entropy-based FilteringYingda Yin, Yingcheng Cai, He Wang, Baoquan ChenCVPR 2022 · 16 citations
Builds on7
- FixMatch: Simplifying Semi-Supervised Learning with Consistency and ConfidenceKihyuk Sohn, David Berthelot, Nicholas Carlini, Zizhao Zhang et al.NeurIPS 2020 · 5,129 citations
- Semi-supervised Keypoint LocalizationOlga Moskvyak, Frédéric Maire, Feras Dayoub, Mahsa BaktashmotlaghICLR 2021 · 17 citations
- Semantic Part Detection via Matching: Learning to Generalize to Novel Viewpoints From Limited Training DataYutong Bai, Qing Liu, Lingxi Xie, Yan Zheng et al.ICCV 2019 · 10 citations
- Compositional Convolutional Neural Networks: A Deep Architecture With Innate Robustness to Partial OcclusionAdam Kortylewski, Ju He, Qing Liu, Alan L. YuilleCVPR 2020
- Robust Object Detection Under Occlusion With Context-Aware CompositionalNetsAngtian Wang, Yihong Sun, Adam Kortylewski, Alan L. YuilleCVPR 2020
Related papers
- Semi-Supervised 3D Hand-Object Poses Estimation With Interactions in TimeShaowei Liu, Hanwen Jiang, Jiarui Xu, Sifei Liu et al.CVPR 2021
- NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose EstimationAngtian Wang, Adam Kortylewski, Alan L. YuilleICLR 2021 · 53 citations
- Weakly-Supervised 3D Human Pose Learning via Multi-View Images in the WildUmar Iqbal, Pavlo Molchanov, Jan KautzCVPR 2020
- ViewNet: Unsupervised Viewpoint Estimation from Conditional GenerationOctave Mariotti, Oisin Mac Aodha, Hakan BilenICCV 2021 · 8 citations
- Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose EstimationPrakhar Kaushik, Aayush Mishra, Adam Kortylewski, Alan L. YuilleICLR 2024 · 10 citations
