FoundationPose: Unified 6D Pose Estimation and Tracking of Novel Objects
Bowen Wen, Wei Yang, Jan Kautz, Stan Birchfield
摘要
We present FoundationPose, a unified foundation model for 6D object pose estimation and tracking, supporting both model-based and model-free setups. Our approach can be instantly applied at test-time to a novel object without finetuning, as long as its CAD model is given, or a small number of reference images are captured. Thanks to the unified framework, the downstream pose estimation modules are the same in both setups, with a neural implicit representation used for efficient novel view synthesis when no CAD model is available. Strong generalizability is achieved via large-scale synthetic training, aided by a large language model (LLM), a novel transformer-based architecture, and contrastive learning formulation. Extensive evaluation on multiple public datasets involving challenging scenarios and objects indicate our unified approach outperforms existing methods specialized for each task by a large margin. In addition, it even achieves comparable results to instance-level methods despite the reduced assumptions. Project page: https://nvlabs.github.io/FoundationPose/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper91
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang 等CVPR 2026 · 被引用 280 次
- SafeVLA: Towards Safety Alignment of Vision-Language-Action Model via Constrained LearningBorong Zhang, Yuhao Zhang, Jiaming Ji, Yingshan Lei 等NeurIPS 2025 · 被引用 84 次
- SoFar: Language-Grounded Orientation Bridges Spatial Reasoning and Object ManipulationZekun Qi, Wenyao Zhang, Yufei Ding, Runpei Dong 等NeurIPS 2025 · 被引用 65 次
- DexMachina: Functional Retargeting for Bimanual Dexterous ManipulationZhao Mandi, Yifan Hou, Dieter Fox, Yashraj Narang 等ICML 2026 · 被引用 55 次
- Neural Assets: 3D-Aware Multi-Object Scene Synthesis with Image Diffusion ModelsZiyi Wu, Yulia Rubanova, Rishabh Kabra, Drew A. Hudson 等NeurIPS 2024 · 被引用 32 次
它引用的顶会 Paper26
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
- Multiview Neural Surface Reconstruction by Disentangling Geometry and AppearanceLior Yariv, Yoni Kasten, Dror Moran, Meirav Galun 等NeurIPS 2020 · 被引用 1,010 次
相关 Paper
- Vision Foundation Model Enables Generalizable Object Pose EstimationKai Chen, Yiyao Ma, Xingyu Lin, Stephen James 等NeurIPS 2024 · 被引用 5 次
- Pos3R: 6D Pose Estimation for Unseen Objects Made EasyWeijian Deng, Dylan Campbell, Chunyi Sun, Jiahao Zhang 等CVPR 2025
- ConceptPose: Training-Free Zero-Shot Object Pose Estimation using Concept VectorsLiming Kuang, Yordanka Velikova, Mahdi Saleh, Jan-Nico Zaech 等CVPR 2026 · 被引用 5 次
- LatentFusion: End-to-End Differentiable Reconstruction and Rendering for Unseen Object Pose EstimationKeunhong Park, Arsalan Mousavian, Yu Xiang, Dieter FoxCVPR 2020
- Any6D: Model-free 6D Pose Estimation of Novel ObjectsTaeyeop Lee, Bowen Wen, Minjun Kang, Gyuree Kang 等CVPR 2025
