3D generation on ImageNet
Ivan Skorokhodov, Aliaksandr Siarohin, Yinghao Xu, Jian Ren, Hsin-Ying Lee, Peter Wonka, Sergey Tulyakov
Abstract
All existing 3D-from-2D generators are designed for well-curated single-category datasets, where all the objects have (approximately) the same scale, 3D location and orientation, and the camera always points to the center of the scene. This makes them inapplicable to diverse, in-the-wild datasets of non-alignable scenes rendered from arbitrary camera poses. In this work, we develop 3D generator with Generic Priors (3DGP): a 3D synthesis framework with more general assumptions about the training data, and show that it scales to very challenging datasets, like ImageNet. Our model is based on three new ideas. First, we incorporate an inaccurate off-the-shelf depth estimator into 3D GAN training via a special depth adaptation module to handle the imprecision. Then, we create a flexible camera model and a regularization strategy for it to learn its distribution parameters during training. Finally, we extend the recent ideas of transferring knowledge from pretrained classifiers into GANs for patch-wise trained models by employing a simple distillation-based technique on top of the discriminator. It achieves more stable training than the existing methods and speeds up the convergence by at least 40%. We explore our model on four datasets: SDIP Dogs 256 2 , SDIP Elephants 256 2 , LSUN Horses 256 2 , and ImageNet 256 2 and demonstrate that 3DGP outperforms the recent state-of-the-art in terms of both texture and geometry quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers35
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long et al.ICLR 2024 · 685 citations
- Text2Tex: Text-driven Texture Synthesis via Diffusion ModelsDave Zhenyu Chen, Yawar Siddiqui, Hsin-Ying Lee, Sergey Tulyakov et al.ICCV 2023 · 262 citations
- DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelYinghao Xu, Hao Tan, Fujun Luan, Sai Bi et al.ICLR 2024 · 234 citations
- DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion PriorJingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang et al.ICLR 2024 · 181 citations
- InfiniCity: Infinite-Scale City SynthesisChieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai et al.ICCV 2023 · 86 citations
Builds on36
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Training Generative Adversarial Networks with Limited DataTero Karras, Miika Aittala, Janne Hellsten, Samuli Laine et al.NeurIPS 2020 · 2,345 citations
- Alias-Free Generative Adversarial NetworksTero Karras, Miika Aittala, Samuli Laine, Erik Härkönen et al.NeurIPS 2021 · 2,126 citations
Related papers
- HoloGAN: Unsupervised Learning of 3D Representations From Natural ImagesThu Nguyen-Phuoc, Chuan Li, Lucas Theis, Christian Richardt et al.ICCV 2019 · 98 citations
- 3DAvatarGAN: Bridging Domains for Personalized Editable AvatarsRameen Abdal, Hsin-Ying Lee, Peihao Zhu, Menglei Chai et al.CVPR 2023
- Learning Generative Models of Textured 3D Meshes from Real-World ImagesDario Pavllo, Jonas Kohler, Thomas Hofmann, Aurélien LucchiICCV 2021 · 57 citations
- 3D-Aware Generative Model for Improved Side-View Image SynthesisKyungmin Jo, Wonjoon Jin, Jaegul Choo, Hyunjoon Lee et al.ICCV 2023 · 5 citations
- Disentangled3D: Learning a 3D Generative Model with Disentangled Geometry and Appearance from Monocular ImagesAyush Tewari, Mallikarjun B. R., Xingang Pan, Ohad Fried et al.CVPR 2022 · 35 citations
