NeMo: Neural Mesh Models of Contrastive Features for Robust 3D Pose Estimation
Angtian Wang, Adam Kortylewski, Alan L. Yuille
Abstract
3D pose estimation is a challenging but important task in computer vision. In this work, we show that standard deep learning approaches to 3D pose estimation are not robust when objects are partially occluded or viewed from a previously unseen pose. Inspired by the robustness of generative vision models to partial occlusion, we propose to integrate deep neural networks with 3D generative representations of objects into a unified neural architecture that we term NeMo. In particular, NeMo learns a generative model of neural feature activations at each vertex on a dense 3D mesh. Using differentiable rendering we estimate the 3D object pose by minimizing the reconstruction error between NeMo and the feature representation of the target image. To avoid local optima in the reconstruction loss, we train the feature extractor to maximize the distance between the individual feature representations on the mesh using contrastive learning. Our extensive experiments on PASCAL3D+, occluded-PASCAL3D+ and ObjectNet3D show that NeMo is much more robust to partial occlusion and unseen pose compared to standard deep networks, while retaining competitive performance on regular data. Interestingly, our experiments also show that NeMo performs reasonably well even when the mesh representation only crudely approximates the true object geometry with a cuboid, hence revealing that the detailed 3D geometry is not needed for accurate 3D pose estimation. The code is publicly available at https://github.com/Angtian/NeMo .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1990b82f-191c-4ff3-98cc-fa26fb763a79Cited by top-tier papers9
- RePOSE: Fast 6D Object Pose Refinement via Deep Texture RenderingShun Iwase, Xingyu Liu, Rawal Khirodkar, Rio Yokota et al.ICCV 2021 · 103 citations
- 3D-Aware Visual Question Answering about Parts, Poses and OcclusionsXingrui Wang, Wufei Ma, Zhuowan Li, Adam Kortylewski et al.NeurIPS 2023 · 27 citations
- Amodal Segmentation through Out-of-Task and Out-of-Distribution Generalization with a Bayesian ModelYihong Sun, Adam Kortylewski, Alan L. YuilleCVPR 2022 · 26 citations
- VoGE: A Differentiable Volume Renderer using Gaussian Ellipsoids for Analysis-by-SynthesisAngtian Wang, Peng Wang, Jian Sun, Adam Kortylewski et al.ICLR 2023 · 4 citations
- Unified Category-Level Object Detection and Pose Estimation from RGB Images Using 3D PrototypesTom Fischer, Xiaojie Zhang, Eddy IlgICCV 2025 · 2 citations
Builds on5
- DPOD: 6D Pose Object Detector and RefinerSergey Zakharov, Ivan Shugurov, Slobodan IlicICCV 2019 · 486 citations
- Compositional Convolutional Neural Networks: A Deep Architecture With Innate Robustness to Partial OcclusionAdam Kortylewski, Ju He, Qing Liu, Alan L. YuilleCVPR 2020
- Robust Object Detection Under Occlusion With Context-Aware CompositionalNetsAngtian Wang, Yihong Sun, Adam Kortylewski, Alan L. YuilleCVPR 2020
- HybridPose: 6D Object Pose Estimation Under Hybrid RepresentationsChen Song, Jiaru Song, Qixing HuangCVPR 2020
- Momentum Contrast for Unsupervised Visual Representation LearningKaiming He, Haoqi Fan, Yuxin Wu, Saining Xie et al.CVPR 2020
Related papers
- 3D-Aware Neural Body Fitting for Occlusion Robust 3D Human Pose EstimationYi Zhang, Pengliang Ji, Angtian Wang, Jieru Mei et al.ICCV 2023 · 44 citations
- Scaling 3D Compositional Models for Robust Classification and Pose EstimationXiaoding Yuan, Guofeng Zhang, Prakhar Kaushik, Artur Jesslen et al.ICCV 2025
- Neural View Synthesis and Matching for Semi-Supervised Few-Shot Learning of 3D PoseAngtian Wang, Shenxiao Mei, Alan L. Yuille, Adam KortylewskiNeurIPS 2021 · 22 citations
- Source-Free and Image-Only Unsupervised Domain Adaptation for Category Level Object Pose EstimationPrakhar Kaushik, Aayush Mishra, Adam Kortylewski, Alan L. YuilleICLR 2024 · 10 citations
- Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature SpaceLeonhard Sommer, Olaf Dünkel, Christian Theobalt, Adam KortylewskiCVPR 2025
