VQ3D: Learning a 3D-Aware Generative Model on ImageNet
Kyle Sargent, Jing Yu Koh, Han Zhang, Huiwen Chang, Charles Herrmann, Pratul P. Srinivasan, Jiajun Wu, Deqing Sun
Abstract
Recent work has shown the possibility of training generative models of 3D content from 2D image collections on small datasets corresponding to a single object class, such as human faces, animal faces, or cars. However, these models struggle on larger, more complex datasets. To model diverse and unconstrained image collections such as ImageNet, we present VQ3D, which introduces a NeRF-based decoder into a two-stage vector-quantized autoencoder. Our Stage 1 allows for the reconstruction of an input image and the ability to change the camera position around the image, and our Stage 2 allows for the generation of new 3D scenes. VQ3D is capable of generating and reconstructing 3D-aware images from the 1000-class ImageNet dataset of 1.2 million training images. We achieve an ImageNet generation FID score of 16.8, compared to 69.8 for the next best baseline method. For video results, please see the project webpage.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c66f1256-eb15-40ee-bc45-205520465103Cited by top-tier papers16
- SyncDreamer: Generating Multiview-consistent Images from a Single-view ImageYuan Liu, Cheng Lin, Zijiao Zeng, Xiaoxiao Long et al.ICLR 2024 · 685 citations
- DreamCraft3D: Hierarchical 3D Generation with Bootstrapped Diffusion PriorJingxiang Sun, Bo Zhang, Ruizhi Shao, Lizhen Wang et al.ICLR 2024 · 181 citations
- 3D-aware Image Generation using 2D Diffusion ModelsJianfeng Xiang, Jiaolong Yang, Binbin Huang, Xin TongICCV 2023 · 82 citations
- Online Clustered CodebookChuanxia Zheng, Andrea VedaldiICCV 2023 · 67 citations
- ZeroNVS: Zero-Shot 360-Degree View Synthesis from a Single ImageKyle Sargent, Zizhang Li, Tanmay Shah, Charles Herrmann et al.CVPR 2024 · 45 citations
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Implicit Neural Representations with Periodic Activation FunctionsVincent Sitzmann, Julien N. P. Martel, Alexander W. Bergman, David B. Lindell et al.NeurIPS 2020 · 4,008 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Class-Partitioned VQ-VAE and Latent Flow Matching for Point Cloud Scene GenerationDasith de Silva Edirimuni, Ajmal Saeed MianAAAI 2026
- G3DR: Generative 3D Reconstruction in ImageNetPradyumna Reddy, Ismail Elezi, Jiankang DengCVPR 2024 · 4 citations
- VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and ReconstructionSinan Du, Jiahao Guo, Bo Li, Shuhao Cui et al.CVPR 2026 · 11 citations
- 3D-aware Blending with Generative NeRFsHyunsu Kim, Gayoung Lee, Yunjey Choi, Jin-Hwa Kim et al.ICCV 2023 · 14 citations
- NeRF-VAE: A Geometry Aware 3D Scene Generative ModelAdam R. Kosiorek, Heiko Strathmann, Daniel Zoran, Pol Moreno et al.ICML 2021 · 167 citations
