EpiGRAF: Rethinking training of 3D GANs
Ivan Skorokhodov, Sergey Tulyakov, Yiqun Wang, Peter Wonka
Abstract
A very recent trend in generative modeling is building 3D-aware generators from 2D image collections. To induce the 3D bias, such models typically rely on volumetric rendering, which is expensive to employ at high resolutions. During the past months, there appeared more than 10 works that address this scaling issue by training a separate 2D decoder to upsample a low-resolution image (or a feature tensor) produced from a pure 3D generator. But this solution comes at a cost: not only does it break multi-view consistency (i.e. shape and texture change when the camera moves), but it also learns the geometry in a low fidelity. In this work, we show that it is possible to obtain a high-resolution 3D generator with SotA image quality by following a completely different route of simply training the model patch-wise. We revisit and improve this optimization scheme in two ways. First, we design a location- and scale-aware discriminator to work on patches of different proportions and spatial positions. Second, we modify the patch sampling strategy based on an annealed beta distribution to stabilize training and accelerate the convergence. The resulted model, named EpiGRAF, is an efficient, high-resolution, pure 3D generator, and we test it on four datasets (two introduced in this work) at and resolutions. It obtains state-of-the-art image quality, high-fidelity geometry and trains faster than the upsampler-based counterparts. Project website: https://universome.github.io/epigraf.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9ba948d9-c9ea-40cb-88ac-149e06212a66Cited by top-tier papers52
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu et al.ICLR 2024 · 408 citations
- DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelYinghao Xu, Hao Tan, Fujun Luan, Sai Bi et al.ICLR 2024 · 234 citations
- Single-Stage Diffusion NeRF: A Unified Approach to 3D Generation and ReconstructionHansheng Chen, Jiatao Gu, Anpei Chen, Wei Tian et al.ICCV 2023 · 213 citations
- GRAM-HD: 3D-Consistent Image Generation at High Resolution with Generative Radiance ManifoldsJianfeng Xiang, Jiaolong Yang, Yu Deng, Xin TongICCV 2023 · 95 citations
- InfiniCity: Infinite-Scale City SynthesisChieh Hubert Lin, Hsin-Ying Lee, Willi Menapace, Menglei Chai et al.ICCV 2023 · 86 citations
Builds on44
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional DomainsMatthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil et al.NeurIPS 2020 · 4,036 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance FieldsJonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman et al.ICCV 2021 · 2,700 citations
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt et al.NeurIPS 2021 · 2,500 citations
Related papers
- What You See is What You GAN: Rendering Every Pixel for High-Fidelity Geometry in 3D GANsAlex Trevithick, Matthew A. Chan, Towaki Takikawa, Umar Iqbal et al.CVPR 2024 · 7 citations
- SinGRAF: Learning a 3D Generative Radiance Field for a Single SceneMinjung Son, Jeong Joon Park, Leonidas J. Guibas, Gordon WetzsteinCVPR 2023
- StyleSDF: High-Resolution 3D-Consistent Image and Geometry GenerationRoy Or-El, Xuan Luo, Mengyi Shan, Eli Shechtman et al.CVPR 2022 · 229 citations
- EVA3D: Compositional 3D Human Generation from 2D Image CollectionsFangzhou Hong, Zhaoxi Chen, Yushi Lan, Liang Pan et al.ICLR 2023 · 35 citations
- En3D: An Enhanced Generative Model for Sculpting 3D Humans from 2D Synthetic DataYifang Men, Biwen Lei, Yuan Yao, Miaomiao Cui et al.CVPR 2024 · 7 citations
