ShapeCrafter: A Recursive Text-Conditioned 3D Shape Generation Model
Rao Fu, Xiao Zhan, Yiwen Chen, Daniel Ritchie, Srinath Sridhar
Abstract
We present ShapeCrafter, a neural network for recursive text-conditioned 3D shape generation. Existing methods that generate text-conditioned 3D shapes consume an entire text prompt to generate a 3D shape in a single step. However, humans tend to describe shapes recursively-we may start with an initial description and progressively add details based on intermediate results. To capture this recursive process, we introduce a method to generate a 3D shape distribution, conditioned on an initial phrase, that gradually evolves as more phrases are added. Since existing datasets are insufficient for training under this approach, we present Text2Shape++, a large dataset of 369K shape-text pairs that supports recursive shape generation. To capture local details that are often used to refine shape descriptions, we build upon vector-quantized deep implicit functions that generate a distribution of high-quality shapes. Results show that our method can generate shapes consistent with text descriptions, and shapes evolve gradually as more phrases are added. Our method supports shape editing, extrapolation, and can enable new applications in human-machine collaboration for creative design.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9eb74cea-7d4f-4fe5-b28a-7b87e6e17232Cited by top-tier papers21
- Instant3D: Fast Text-to-3D with Sparse-view Generation and Large Reconstruction ModelJiahao Li, Hao Tan, Kai Zhang, Zexiang Xu et al.ICLR 2024 · 408 citations
- Locally Attentional SDF Diffusion for Controllable 3D Shape GenerationXin-Yang Zheng, Hao Pan, Peng-Shuai Wang, Xin Tong et al.SIGGRAPH 2023 · 122 citations
- SKED: Sketch-guided Text-based 3D EditingAryan Mikaeili, Or Perel, Mehdi Safaee, Daniel Cohen-Or et al.ICCV 2023 · 83 citations
- VPP: Efficient Conditional 3D Generation via Voxel-Point Progressive RepresentationZekun Qi, Muzhou Yu, Runpei Dong, Kaisheng MaNeurIPS 2023 · 22 citations
- Image Content Generation with Causal ReasoningXiaochuan Li, Baoyu Fan, Runze Zhang, Liang Jin et al.AAAI 2024 · 13 citations
Builds on14
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Zero-Shot Text-Guided Object Generation with Dream FieldsAjay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel et al.CVPR 2022 · 361 citations
- CLIP-NeRF: Text-and-Image Driven Manipulation of Neural Radiance FieldsCan Wang, Menglei Chai, Mingming He, Dongdong Chen et al.CVPR 2022 · 313 citations
Related papers
- ShapeFormer: Transformer-based Shape Completion via Sparse RepresentationXingguang Yan, Liqiang Lin, Niloy J. Mitra, Dani Lischinski et al.CVPR 2022 · 124 citations
- ShapeCraft: LLM Agents for Structured, Textured and Interactive 3D ModelingShuyuan Zhang, Chenhan Jiang, Zuoou Li, Jiankang DengNeurIPS 2025 · 6 citations
- ShapeScaffolder: Structure-Aware 3D Shape Generation from TextXi Tian, Yong-Liang Yang, Qi WuICCV 2023 · 14 citations
- ShapeWalk: Compositional Shape Editing Through Language-Guided ChainsHabib Slim, Mohamed ElhoseinyCVPR 2024 · 4 citations
- SPAGHETTI: editing implicit shapes through part aware generationAmir Hertz, Or Perel, Raja Giryes, Olga Sorkine-Hornung et al.SIGGRAPH 2022 · 59 citations
