CUPID: Generative 3D Reconstruction via Joint Object and Pose Modeling
Binbin Huang, Haobin Duan, Yiqun Zhao, Zibo Zhao, Yi Ma, Shenghua Gao
Abstract
We introduce Cupid, a generative 3D reconstruction framework that jointly models the full distribution over both canonical objects and camera poses. Our two-stage flow-based model first generates a coarse 3D structure and 2D-3D correspondences to estimate the camera pose robustly. Conditioned on this pose, a refinement stage injects pixel-aligned image features directly into the generative process, marrying the rich prior of a generative model with the geometric fidelity of reconstruction. This strategy achieves exceptional faithfulness, outperforming state-of-the-art reconstruction methods by over 3 dB PSNR and 10% in Chamfer Distance. As a unified generative model that decouples the object and camera pose, Cupid naturally extends to multi-view and scene-level reconstruction tasks without requiring post-hoc optimization or fine-tuning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bfe66f82-0b2c-47f2-9762-8b1e448f4bdaCited by top-tier papers3
- LATTICE: Democratize High-Fidelity 3D Generation at ScaleZeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu et al.CVPR 2026 · 46 citations
- Mix3R: Mixing Feed-forward Reconstruction and Generative 3D Priors for Joint Multi-view Aligned 3D Reconstruction and Pose EstimationSiyou Lin, Zhou Xue, Hongwen Zhang, Liang An et al.SIGGRAPH 2026
- AniGen: Unified S3 Fields for Animatable 3D Asset GenerationYihua Huang, Zi-Xin Zou, Yuting He, Chirui Chang et al.SIGGRAPH 2026
Builds on47
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
- Depth Anything V2Lihe Yang, Bingyi Kang, Zilong Huang, Zhen Zhao et al.NeurIPS 2024 · 2,305 citations
Related papers
- Pixal3D: Pixel-Aligned 3D Generation from ImagesDong-Yang Li, Wang Zhao, Yuxin Chen, Wenbo Hu et al.SIGGRAPH 2026 · 1 citation
- SyncHuman: Synchronizing 2D and 3D Generative Models for Single-view Human ReconstructionWenyue Chen, Peng Li, Wangguandong Zheng, Chengfeng Zhao et al.NeurIPS 2025 · 8 citations
- JRM: Joint Reconstruction Model for Multiple Objects without AlignmentQirui Wu, Mohd Yawar Nihal Siddiqui, Duncan Frost, Samir Aroudj et al.CVPR 2026 · 2 citations
- TokenSplat: Token-aligned 3D Gaussian Splatting for Feed-forward Pose-free ReconstructionYihui Li, Chengxin Lv, Zichen Tang, Hongyu Yang et al.CVPR 2026 · 13 citations
- GAUDI: A Neural Architect for Immersive 3D Scene GenerationMiguel Ángel Bautista, Pengsheng Guo, Samira Abnar, Walter Talbott et al.NeurIPS 2022 · 170 citations
