Make-A-Shape: a Ten-Million-scale 3D Shape Model
Ka-Hei Hui, Aditya Sanghi, Arianna Rampini, Kamal Rahimi Malekshan, Zhengzhe Liu, Hooman Shayani, Chi-Wing Fu
Abstract
Significant progress has been made in training large generative models for natural language and images. Yet, the advancement of 3D generative models is hindered by their substantial resource demands for training, along with inefficient, non-compact, and less expressive representations. This paper introduces Make-A-Shape, a new 3D generative model designed for efficient training on a vast scale, capable of utilizing 10 millions publicly-available shapes. Technical-wise, we first innovate a wavelet-tree representation to compactly encode shapes by formulating the subband coefficient filtering scheme to efficiently exploit coefficient relations. We then make the representation generatable by a diffusion model by devising the subband coefficients packing scheme to layout the representation in a low-resolution grid. Further, we derive the subband adaptive training strategy to train our model to effectively learn to generate coarse and detail wavelet coefficients. Last, we extend our framework to be controlled by additional input conditions to enable it to generate shapes from assorted modalities, e.g., single/multi-view images, point clouds, and low-resolution voxels. In our extensive set of experiments, we demonstrate various applications, such as unconditional generation, shape completion, and conditional generation on a wide range of modalities. Our approach not only surpasses the state of the art in delivering high-quality results but also efficiently generates shapes within a few seconds, often achieving this in just 2 seconds for most conditions. Our source code is available at https://github.com/AutodeskAILab/Make-a-Shape.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers11
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang et al.NeurIPS 2024 · 251 citations
- MeshFormer : High-Quality Mesh Generation with 3D-Guided Reconstruction ModelMinghua Liu, Chong Zeng, Xinyue Wei, Ruoxi Shi et al.NeurIPS 2024 · 73 citations
- X-Ray: A Sequential 3D Representation For GenerationTao Hu, Wenhang Ge, Yuyang Zhao, Gim Hee LeeNeurIPS 2024 · 11 citations
- CoFie: Learning Compact Neural Surface Representations with Coordinate FieldsHanwen Jiang, Haitao Yang, Georgios Pavlakos, Qixing HuangNeurIPS 2024 · 6 citations
- PrimitiveAnything: Human-Crafted 3D Primitive Assembly Generation with Auto-Regressive transformerJingwen Ye, Yuze He, Yanning Zhou, Yiqin Zhu et al.SIGGRAPH 2025 · 5 citations
Builds on51
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
Related papers
- OctGPT: Octree-based Multiscale Autoregressive Models for 3D Shape GenerationSi-Tong Wei, Rui-Huan Wang, Chuan-Zhi Zhou, Baoquan Chen et al.SIGGRAPH 2025 · 11 citations
- Efficient Autoregressive Shape Generation Via Octree-Based Adaptive TokenizationKangle Deng, Hsueh-Ti Derek Liu, Yiheng Zhu, Xiaoxia Sun et al.ICCV 2025 · 4 citations
- VPP: Efficient Conditional 3D Generation via Voxel-Point Progressive RepresentationZekun Qi, Muzhou Yu, Runpei Dong, Kaisheng MaNeurIPS 2023 · 22 citations
- 3DShape2VecSet: A 3D Shape Representation for Neural Fields and Generative Diffusion ModelsBiao Zhang, Jiapeng Tang, Matthias Nießner, Peter WonkaSIGGRAPH 2023 · 172 citations
- 3D Shape Generation and Completion through Point-Voxel DiffusionLinqi Zhou, Yilun Du, Jiajun WuICCV 2021 · 681 citations
