ShapeScaffolder: Structure-Aware 3D Shape Generation from Text
Xi Tian, Yong-Liang Yang, Qi Wu
Abstract
We present ShapeScaffolder, a structure-based neural network for generating colored 3D shapes based on text input. The approach, similar to providing scaffolds as internal structural supports and adding more details to them, aims to capture finer text-shape connections and improve the quality of generated shapes. Traditional text-to-shape methods often generate 3D shapes as a whole. However, humans tend to understand both shape and text as being structure-based. For example, a table is interpreted as being composed of legs, a seat, and a back; similarly, texts possess inherent linguistic structures that can be analyzed as dependency graphs, depicting the relationships between entities within the text. We believe structure-aware shape generation can bring finer text-shape connections and improve shape generation quality. However, the lack of explicit shape structure and the high freedom of text structure make cross-modality learning challenging. To address these challenges, we first build the structured shape implicit fields in an unsupervised manner. We then propose the part-level attention mechanism between shape parts and textual graph nodes to align the two modalities at the structural level. Finally, we employ a shape refiner to add further detail to the predicted structure, yielding the final results. Extensive experimentation demonstrates that our approaches outperform state-of-the-art methods in terms of both shape fidelity and shape-text matching. Our methods also allow for part-level manipulation and improved part-level completeness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- Tera: Rethinking Text-Guided Realistic 3D Avatar GenerationYanwen Wang, Yiyu Zhuang, Jiawei Zhang, Li Wang et al.ICCV 2025 · 2 citations
- HyperSDFusion: Bridging Hierarchical Structures in Language and Geometry for Enhanced 3D Text2Shape GenerationZhiying Leng, Tolga Birdal, Xiaohui Liang, Federico TombariCVPR 2024
- Rapid 3D Model Generation with Intuitive 3D InputTianrun Chen, Chaotao Ding, Shangzhan Zhang, Chunan Yu et al.CVPR 2024
- Order Matters: 3D Shape Generation from Sequential VR SketchesYizi Chen, Sidi Wu, Tianyi Xiao, Nina Wiedemann et al.CVPR 2026
Builds on16
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 463 citations
- Zero-Shot Text-Guided Object Generation with Dream FieldsAjay Jain, Ben Mildenhall, Jonathan T. Barron, Pieter Abbeel et al.CVPR 2022 · 361 citations
- CLIP-Forge: Towards Zero-Shot Text-to-Shape GenerationAditya Sanghi, Hang Chu, Joseph G. Lambourne, Ye Wang et al.CVPR 2022 · 206 citations
- Specifying Object Attributes and Relations in Interactive Scene GenerationOron Ashual, Lior WolfICCV 2019 · 190 citations
Related papers
- Towards Implicit Text-Guided 3D Shape GenerationZhengzhe Liu, Yi Wang, Xiaojuan Qi, Chi-Wing FuCVPR 2022 · 59 citations
- ShapeCaptioner: Generative Caption Network for 3D Shapes by Learning a Mapping from Parts Detected in Multiple Views to SentencesZhizhong Han, Chao Chen, Yu-Shen Liu, Matthias ZwickerACM MM 2020 · 38 citations
- ShapeCrafter: A Recursive Text-Conditioned 3D Shape Generation ModelRao Fu, Xiao Zhan, Yiwen Chen, Daniel Ritchie et al.NeurIPS 2022 · 98 citations
- Michelangelo: Conditional 3D Shape Generation based on Shape-Image-Text Aligned Latent RepresentationZibo Zhao, Wen Liu, Xin Chen, Xianfang Zeng et al.NeurIPS 2023 · 279 citations
- DensiCrafter: Physically-Constrained Generation and Fabrication of Self-Supporting Hollow StructuresShengqi Dang, Fu Chai, Jiaxin Li, Chao Yuan et al.AAAI 2026 · 1 citation
