ConRad: Image Constrained Radiance Fields for 3D Generation from a Single Image
Senthil Purushwalkam, Nikhil Naik
摘要
We present a novel method for reconstructing 3D objects from a single RGB image. Our method leverages the latest image generation models to infer the hidden 3D structure while remaining faithful to the input image. While existing methods [1, 2] obtain impressive results in generating 3D models from text prompts, they do not provide an easy approach for conditioning on input RGB data. Naïve extensions of these methods often lead to improper alignment in appearance between the input image and the 3D reconstructions. We address these challenges by introducing Image Constrained Radiance Fields (ConRad), a novel variant of neural radiance fields. ConRad is an efficient 3D representation that explicitly captures the appearance of an input image in one viewpoint. We propose a training algorithm that leverages the single RGB image in conjunction with pretrained Diffusion Models to optimize the parameters of a ConRad representation. Extensive experiments show that ConRad representations can simplify preservation of image details while producing a realistic 3D reconstruction. Compared to existing state-of-the-art baselines, we show that our 3D reconstructions remain more faithful to the input and produce more consistent 3D models while demonstrating significantly improved quantitative performance on a ShapeNet object benchmark.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Free3D: Consistent Novel View Synthesis Without 3D RepresentationChuanxia Zheng, Andrea VedaldiCVPR 2024 · 被引用 28 次
- MVD^2: Efficient Multiview 3D Reconstruction for Multiview DiffusionXin-Yang Zheng, Hao Pan, Yu-Xiao Guo, Xin Tong 等SIGGRAPH 2024 · 被引用 12 次
- LucidDreamer: Towards High-Fidelity Text-to-3D Generation via Interval Score MatchingYixun Liang, Xin Yang, Jiantao Lin, Haodong Li 等CVPR 2024
它引用的顶会 Paper23
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language ModelsJunnan Li, Dongxu Li, Silvio Savarese, Steven C. H. HoiICML 2023 · 被引用 7,873 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
相关 Paper
- Make-It-3D: High-Fidelity 3D Creation from A Single Image with Diffusion PriorJunshu Tang, Tengfei Wang, Bo Zhang, Ting Zhang 等ICCV 2023 · 被引用 405 次
- One-2-3-45: Any Single Image to 3D Mesh in 45 Seconds without Per-Shape OptimizationMinghua Liu, Chao Xu, Haian Jin, Linghao Chen 等NeurIPS 2023 · 被引用 755 次
- LOLNeRF: Learn from One LookDaniel Rebain, Mark J. Matthews, Kwang Moo Yi, Dmitry Lagun 等CVPR 2022
- LRM: Large Reconstruction Model for Single Image to 3DYicong Hong, Kai Zhang, Jiuxiang Gu, Sai Bi 等ICLR 2024 · 被引用 813 次
- RenderDiffusion: Image Diffusion for 3D Reconstruction, Inpainting and GenerationTitas Anciukevicius, Zexiang Xu, Matthew Fisher, Paul Henderson 等CVPR 2023
