RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose Estimation
Junwen Huang, Shishir Reddy Vutukur, Peter KT Yu, Nassir Navab, Slobodan Ilic, Benjamin Busam
摘要
Typical template-based object pose pipelines estimate the pose by retrieving the closest matching template and aligning it with the observed image. However, failure to retrieve the correct template often leads to inaccurate pose predictions. To address this, we reformulate template-based object pose estimation as a ray alignment problem, where the viewing directions from multiple posed template images are learned to align with a non-posed query image. Inspired by recent progress in diffusion-based camera pose estimation, we embed this formulation into a diffusion transformer architecture that aligns a query image with a set of posed templates. We reparameterize object rotation using object-centered camera rays and model object translation by extending scale-invariant translation estimation to dense translation offsets. Our model leverages geometric priors from the templates to guide accurate query pose inference. A coarse-to-fine training strategy based on narrowed template sampling improves performance without modifying the network architecture. Extensive experiments across multiple benchmark datasets show competitive results of our method compared to state-of-the-art approaches in unseen object pose estimation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- 3D-Object Perception Transformer (3PT)Agastya Kalra, Tim Salzmann, Guy Stoppi, Dmitrii Marin 等CVPR 2026 · 被引用 1 次
- PoseGAM: Robust Unseen Object Pose Estimation via Geometry-Aware Multi-View ReasoningJianqi Chen, Biao Zhang, Xiangjun Tang, Peter WonkaCVPR 2026
它引用的顶会 Paper27
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- MVDream: Multi-view Diffusion for 3D GenerationYichun Shi, Peng Wang, Jianglong Ye, Long Mai 等ICLR 2024 · 被引用 973 次
- Perceiver IO: A General Architecture for Structured Inputs & OutputsAndrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch 等ICLR 2022 · 被引用 797 次
- MVDiffusion: Enabling Holistic Multi-view Image Generation with Correspondence-Aware DiffusionShitao Tang, Fuyang Zhang, Jiacheng Chen, Peng Wang 等NeurIPS 2023 · 被引用 249 次
相关 Paper
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang 等ICLR 2024 · 被引用 126 次
- DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint DiffusionQitao Zhao, Amy Lin, Jeff Tan, Jason Y. Zhang 等CVPR 2025
- FAR: Flexible, Accurate and Robust 6DoF Relative Camera Pose EstimationChris Rockwell, Nilesh Kulkarni, Linyi Jin, Jeong Joon Park 等CVPR 2024
- Composing People Together: Iterative Pose-Image Generation for Multi-Person Interaction ScenesWenxuan Peng, Bharath Hariharan, Hadar Averbuch-ElorSIGGRAPH 2026
- Co-op: Correspondence-based Novel Object Pose EstimationSungphill Moon, Hyeontae Son, Dongcheol Hur, Sangwook KimCVPR 2025
