Unsupervised Learning of Category-Level 3D Pose from Object-Centric Videos
Leonhard Sommer, Artur Jesslen, Eddy Ilg, Adam Kortylewski
摘要
Category-level 3D pose estimation is a fundamentally important problem in computer vision and robotics, e.g. for embodied agents or to train 3D generative models. However, so far methods that estimate the category-level object pose require either large amounts of human annotations, CAD models or input from RGB-D sensors. In contrast, we tackle the problem of learning to estimate the category-level 3D pose only from casually taken object-centric videos without human supervision. We propose a two-step pipeline: First, we introduce a multi-view alignment procedure that determines canonical camera poses across videos with a novel and robust cyclic distance formulation for geometric and appearance matching using reconstructed coarse meshes and DINOv2 features. In a second step, the canonical poses and reconstructed meshes enable us to train a model for 3D pose estimation from a single image. In particular, our model learns to estimate dense correspondences between images and a prototypical 3D template by predicting, for each pixel in a 2D image, a feature vector of the corresponding vertex in the template mesh. We demonstrate that our method outperforms all baselines at the unsupervised alignment of object-centric videos by a large margin and provides faithful and robust predictions in-the-wild on the Pascal3D+ and ObjectNet3D datasets.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Pose Prior Learner: Unsupervised Categorical Prior Learning for Pose EstimationZiyu Wang, Shuangpeng Han, Mengmi ZhangICLR 2026 · 被引用 3 次
- One-shot 3D Object Canonicalization based on Geometric and Semantic ConsistencyLi Jin, Yujie Wang, Wenzheng Chen, Qiyu Dai 等CVPR 2025
- Common3D: Self-Supervised Learning of 3D Morphable Models for Common Objects in Neural Feature SpaceLeonhard Sommer, Olaf Dünkel, Christian Theobalt, Adam KortylewskiCVPR 2025
它引用的顶会 Paper10
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
- SGPA: Structure-Guided Prior Adaptation for Category-Level 6D Object Pose EstimationKai Chen, Qi DouICCV 2021 · 被引用 183 次
- Continuous Surface EmbeddingsNatalia Neverova, David Novotný, Marc Szafraniec, Vasil Khalidov 等NeurIPS 2020 · 被引用 116 次
- Category-Level 6D Object Pose Estimation in the Wild: A Semi-Supervised Learning Approach and A New DatasetYanjie Ze, Xiaolong WangNeurIPS 2022 · 被引用 104 次
相关 Paper
- Self-Supervised Geometric Correspondence for Category-Level 6D Object Pose Estimation in the WildKaifeng Zhang, Yang Fu, Shubhankar Borse, Hong Cai 等ICLR 2023 · 被引用 8 次
- SecondPose: SE(3)-Consistent Dual-Stream Feature Fusion for Category-Level Pose EstimationYamei Chen, Yan Di, Guangyao Zhai, Fabian Manhardt 等CVPR 2024
- Leveraging SE(3) Equivariance for Self-supervised Category-Level Object Pose Estimation from Point CloudsXiaolong Li, Yijia Weng, Li Yi, Leonidas J. Guibas 等NeurIPS 2021 · 被引用 61 次
- Novel Object Viewpoint Estimation Through Reconstruction AlignmentMohamed El Banani, Jason J. Corso, David F. FouheyCVPR 2020
- Category-Level Articulated Object Pose EstimationXiaolong Li, He Wang, Li Yi, Leonidas J. Guibas 等CVPR 2020
