PF-LRM: Pose-Free Large Reconstruction Model for Joint Pose and Shape Prediction
Peng Wang, Hao Tan, Sai Bi, Yinghao Xu, Fujun Luan, Kalyan Sunkavalli, Wenping Wang, Zexiang Xu, Kai Zhang
摘要
We propose a Pose-Free Large Reconstruction Model (PF-LRM) for reconstructing a 3D object from a few unposed images even with little visual overlap, while simultaneously estimating the relative camera poses in 1.3 seconds on a single A100 GPU. PF-LRM is a highly scalable method utilizing the self-attention blocks to exchange information between 3D object tokens and 2D image tokens; we predict a coarse point cloud for each view, and then use a differentiable Perspective-n-Point (PnP) solver to obtain camera poses. When trained on a huge amount of multi-view posed data of 1M objects, PF-LRM shows strong cross-dataset generalization ability, and outperforms baseline methods by a large margin in terms of pose prediction accuracy and 3D reconstruction quality on various unseen evaluation datasets. We also demonstrate our model's applicability in downstream text/image-to-3D task with fast feed-forward inference. Our project website is at: https://totoro97.github.io/pf-lrm .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper85
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
- CLAY: A Controllable Large-scale Generative Model for Creating High-quality 3D AssetsLongwen Zhang, Ziyu Wang, Qixuan Zhang, Qiwei Qiu 等SIGGRAPH 2024 · 被引用 148 次
- Era3D: High-Resolution Multiview Diffusion using Efficient Row-wise AttentionPeng Li, Yuan Liu, Xiaoxiao Long, Feihu Zhang 等NeurIPS 2024 · 被引用 132 次
- Cameras as Rays: Pose Estimation via Ray DiffusionJason Y. Zhang, Amy Lin, Moneish Kumar, Tzu-Hsuan Yang 等ICLR 2024 · 被引用 126 次
- MeshXL: Neural Coordinate Field for Generative 3D Foundation ModelsSijin Chen, Xin Chen, Anqi Pang, Xianfang Zeng 等NeurIPS 2024 · 被引用 125 次
它引用的顶会 Paper37
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
相关 Paper
- FreeSplatter: Pose-free Gaussian Splatting for Sparse-view 3D ReconstructionJiale Xu, Shenghua Gao, Ying ShanICCV 2025 · 被引用 8 次
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi 等ICLR 2024 · 被引用 813 次
- GeoLRM: Geometry-Aware Large Reconstruction Model for High-Quality 3D Gaussian GenerationChubin Zhang, Hongliang Song, Yi Wei, Chen Yu 等NeurIPS 2024 · 被引用 40 次
- TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free ReconstructionYihui Li, Chengxin Lv, Zichen Tang, Hongyu Yang 等CVPR 2026 · 被引用 13 次
- PE3R: Perception-Efficient 3D ReconstructionJie Hu, Shizun Wang, Xinchao WangCVPR 2026 · 被引用 9 次
