Cameras as Rays: Pose Estimation via Ray Diffusion
Jason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang, Deva Ramanan, Shubham Tulsiani
摘要
Estimating camera poses is a fundamental task for 3D reconstruction and remains challenging given sparsely sampled views (<10). In contrast to existing approaches that pursue top-down prediction of global parametrizations of camera extrinsics, we propose a distributed representation of camera pose that treats a camera as a bundle of rays. This representation allows for a tight coupling with spatial image features improving pose precision. We observe that this representation is naturally suited for set-level transformers and develop a regression-based approach that maps image patches to corresponding rays. To capture the inherent uncertainties in sparse-view pose inference, we adapt this approach to learn a denoising diffusion model which allows us to sample plausible modes while improving performance. Our proposed methods, both regression-and diffusion-based, demonstrate state-of-the-art performance on camera pose estimation on CO3D while generalizing to unseen object categories and in-the-wild captures. Ray Diffusion Timesteps Recovered Cameras Images * denotes equal contribution. Project Page: https://jasonyzhang.com/RayDiffusion .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper64
- Cameras as Relative Positional EncodingRuilong Li, Brent Yi, Junchen Liu, Hang Gao 等NeurIPS 2025 · 被引用 113 次
- Dens3R: A Foundation Model for 3D Geometry PredictionXianze Fang, Jingnan Gao, Zhe Wang, Zhuo Chen 等ICLR 2026 · 被引用 45 次
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 被引用 22 次
- PlayerOne: Egocentric World SimulatorYuanpeng Tu, Hao Luo, Xi Chen, Xiang Bai 等NeurIPS 2025 · 被引用 20 次
- Thinking with Camera: A Unified Multimodal Model for Camera-Centric Understanding and GenerationKang Liao, Size Wu, Zhonghua Wu, Linyi Jin 等ICLR 2026 · 被引用 19 次
它引用的顶会 Paper19
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Common Objects in 3D: Large-Scale Learning and Evaluation of Real-life 3D Category ReconstructionJeremy Reizenstein, Roman Shapovalov, Philipp Henzler, Luca Sbordone 等ICCV 2021 · 被引用 686 次
- Light Field Networks: Neural Scene Representations with Single-Evaluation RenderingVincent Sitzmann, Semon Rezchikov, Bill Freeman, Josh Tenenbaum 等NeurIPS 2021 · 被引用 426 次
相关 Paper
- RayPose: Ray Bundling Diffusion for Template Views in Unseen 6D Object Pose EstimationJunwen Huang, Shishir Reddy Vutukur, Peter KT Yu, Nassir Navab 等ICCV 2025 · 被引用 1 次
- GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray DiffusionLi-Heng Chen, Zi-Xin Zou, Chang Liu, Tianjiao Jing 等ICCV 2025
- DiffusionSfM: Predicting Structure and Motion via Ray Origin and Endpoint DiffusionQitao Zhao, Amy Lin, Jeff Tan, Jason Y. Zhang 等CVPR 2025
- PoseDiffusion: Solving Pose Estimation via Diffusion-aided Bundle AdjustmentJianyuan Wang, Christian Rupprecht, David NovotnýICCV 2023 · 被引用 158 次
- AlignDiff: Learning Physically-Grounded Camera Alignment via DiffusionLiuyue Xie, Jiancong Guo, Ozan Cakmakci, Andre Araujo 等ICCV 2025
