PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion Transformers
Yuchen Lin, Chenguo Lin, Panwang Pan, Honglei Yan, Yiqiang Feng, Yadong Mu, Katerina Fragkiadaki
摘要
We introduce PartCrafter, the first structured 3D generative model that jointly synthesizes multiple semantically meaningful and geometrically distinct 3D meshes from a single RGB image. Unlike existing methods that either produce monolithic 3D shapes or follow two-stage pipelines, i.e., first segmenting an image and then reconstructing each segment, PartCrafter adopts a unified, compositional generation architecture that does not rely on pre-segmented inputs. Conditioned on a single image, it simultaneously denoises multiple 3D parts, enabling end-to-end part-aware generation of both individual objects and complex multi-object scenes. PartCrafter builds upon a pretrained 3D mesh diffusion transformer (DiT) trained on whole objects, inheriting the pretrained weights, encoder, and decoder, and introduces two key innovations: (1) A compositional latent space, where each 3D part is represented by a set of disentangled latent tokens; (2) A hierarchical attention mechanism that enables structured information flow both within individual parts and across all parts, ensuring global coherence while preserving part-level detail during generation. To support part-level supervision, we curate a new dataset by mining part-level annotations from large-scale 3D object datasets. Experiments show that PartCrafter outperforms existing approaches in generating decomposable 3D meshes, including parts that are not directly visible in input images, demonstrating the strength of part-aware generative priors for 3D understanding and synthesis. Code and training data will be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper25
- SceneMaker: Open-set 3D Scene Generation with Decoupled De-occlusion and Pose Estimation ModelYukai Shi, Weiyu Li, Zihao Wang, Hongyang Li 等CVPR 2026 · 被引用 17 次
- ART: Articulated Reconstruction TransformerZizhang Li, Cheng Zhang, Zhengqin Li, Henry Howard-Jenkins 等CVPR 2026 · 被引用 12 次
- 3D-Fixer: Coarse-to-Fine In-place Completion for 3D Scenes from a Single ImageZe-Xin Yin, Liu Liu, Xinjie wang, Wei Sui 等CVPR 2026 · 被引用 10 次
- UniPart: Part-Level 3D Generation with Unified 3D Geom-Seg LatentsXufan He, Yushuang Wu, Xiaoyang Guo, Chongjie Ye 等CVPR 2026 · 被引用 9 次
- MeshMosaic: Scaling Artist Mesh Generation via Local-to-Global AssemblyRui Xu, Tianyang Xue, Qiujie Dong, Le Wan 等CVPR 2026 · 被引用 8 次
它引用的顶会 Paper54
- Segment AnythingAlexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao 等ICCV 2023 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- NeuS: Learning Neural Implicit Surfaces by Volume Rendering for Multi-view ReconstructionPeng Wang, Lingjie Liu, Yuan Liu, Christian Theobalt 等NeurIPS 2021 · 被引用 2,500 次
- LION: Latent Point Diffusion Models for 3D Shape GenerationXiaohui Zeng, Arash Vahdat, Francis Williams, Zan Gojcic 等NeurIPS 2022 · 被引用 752 次
相关 Paper
- PartDiffuser: Part-wise 3D Mesh Generation via Discrete DiffusionYichen Yang, Hong Li, Haodong Zhu, Linin Yang 等CVPR 2026 · 被引用 4 次
- HumanCrafter: Synergizing Generalizable Human Reconstruction and Semantic 3D SegmentationPanwang Pan, Tingting Shen, Chenxin Li, Yunlong Lin 等NeurIPS 2025
- Particulate: Feed-Forward 3D Object ArticulationRuining Li, Yuxin Yao, Chuanxia Zheng, Christian Rupprecht 等CVPR 2026 · 被引用 22 次
- Efficient Part-level 3D Object Generation via Dual Volume PackingJiaxiang Tang, Ruijie Lu, Max Li, Zekun Hao 等NeurIPS 2025 · 被引用 53 次
- Quartet of Diffusions: Structure-Aware Point Cloud Generation through Part and Symmetry GuidanceChenliang Zhou, Fangcheng Zhong, Weihao Xia, Albert Miao 等ICLR 2026
