Large-Vocabulary 3D Diffusion Model with Transformer
Ziang Cao, Fangzhou Hong, Tong Wu, Liang Pan, Ziwei Liu
摘要
Creating diverse and high-quality 3D assets with an automatic generative model is highly desirable. Despite extensive efforts on 3D generation, most existing works focus on the generation of a single category or a few categories. In this paper, we introduce a diffusion-based feed-forward framework for synthesizing massive categories of real-world 3D objects with a single generative model. Notably, there are three major challenges for this large-vocabulary 3D generation: a) the need for expressive yet efficient 3D representation; b) large diversity in geometry and texture across categories; c) complexity in the appearances of real-world objects. To this end, we propose a novel triplane-based 3D-aware Diffusion model with TransFormer, DiffTF, for handling challenges via three aspects. 1) Considering efficiency and robustness, we adopt a revised triplane representation and improve the fitting speed and accuracy. 2) To handle the drastic variations in geometry and texture, we regard the features of all 3D objects as a combination of generalized 3D knowledge and specialized 3D features. To extract generalized 3D knowledge from diverse categories, we propose a novel 3D-aware transformer with shared cross-plane attention. It learns the cross-plane relations across different planes and aggregates the generalized 3D knowledge with specialized 3D features. 3) In addition, we devise the 3D-aware encoder/decoder to enhance the generalized 3D knowledge in the encoded triplanes for handling categories with complex appearances. Extensive experiments on ShapeNet and OmniObject3D (over 200 diverse real-world categories) convincingly demonstrate that a single DiffTF model achieves state-of-the-art large-vocabulary 3D object generation performance with large diversity, rich semantics, and high quality. Our project page: https:// ziangcao0312.github.io/difftf_pages/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper24
- Learning-to-Cache: Accelerating Diffusion Transformer via Layer CachingXinyin Ma, Gongfan Fang, Michael Bi Mi, Xinchao WangNeurIPS 2024 · 被引用 167 次
- PartCrafter: Structured 3D Mesh Generation via Compositional Latent Diffusion TransformersYuchen Lin, Chenguo Lin, Panwang Pan, Honglei Yan 等NeurIPS 2025 · 被引用 89 次
- DiffGS: Functional Gaussian Splatting DiffusionJunsheng Zhou, Weiqi Zhang, Yu-Shen LiuNeurIPS 2024 · 被引用 69 次
- Efficient Part-level 3D Object Generation via Dual Volume PackingJiaxiang Tang, Ruijie Lu, Max Li, Zekun Hao 等NeurIPS 2025 · 被引用 53 次
- GaussianCube: A Structured and Explicit Radiance Representation for 3D Generative ModelingBowen Zhang, Yiji Cheng, Jiaolong Yang, Chunyu Wang 等NeurIPS 2024 · 被引用 49 次
它引用的顶会 Paper24
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Deformable DETR: Deformable Transformers for End-to-End Object DetectionXizhou Zhu, Weijie Su, Lewei Lu, Bin Li 等ICLR 2021 · 被引用 7,353 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- DMV3D: Denoising Multi-view Diffusion Using 3D Large Reconstruction ModelYinghao Xu, Hao Tan, Fujun Luan, Sai Bi 等ICLR 2024 · 被引用 234 次
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
- 3D Neural Field Generation Using Triplane DiffusionJ. Ryan Shue, Eric Ryan Chan, Ryan Po, Zachary Ankner 等CVPR 2023
- DIRECT-3D: Learning Direct Text-to-3D Generation on Massive Noisy 3D DataQihao Liu, Yi Zhang, Song Bai, Adam Kortylewski 等CVPR 2024 · 被引用 4 次
- TAR3D: Creating High-Quality 3D Assets Via Next-Part PredictionXuying Zhang, Yutong Liu, Yangguang Li, Renrui Zhang 等ICCV 2025 · 被引用 4 次
