DiT-3D: Exploring Plain Diffusion Transformers for 3D Shape Generation
Shentong Mo, Enze Xie, Ruihang Chu, Lanqing Hong, Matthias Nießner, Zhenguo Li
Abstract
Recent Diffusion Transformers (e.g. DiT [1] ) have demonstrated their powerful effectiveness in generating high-quality 2D images. However, it is still being determined whether the Transformer architecture performs equally well in 3D shape generation, as previous 3D diffusion methods mostly adopted the U-Net architecture. To bridge this gap, we propose a novel Diffusion Transformer for 3D shape generation, namely DiT-3D, which can directly operate the denoising process on voxelized point clouds using plain Transformers. Compared to existing U-Net approaches, our DiT-3D is more scalable in model size and produces much higher quality generations. Specifically, the DiT-3D adopts the design philosophy of DiT [1] but modifies it by incorporating 3D positional and patch embeddings to adaptively aggregate input from voxelized point clouds. To reduce the computational cost of self-attention in 3D shape generation, we incorporate 3D window attention into Transformer blocks, as the increased 3D token length resulting from the additional dimension of voxels can lead to high computation. Finally, linear and devoxelization layers are used to predict the denoised point clouds. In addition, our transformer architecture supports efficient fine-tuning from 2D to 3D, where the pre-trained DiT-2D checkpoint on ImageNet can significantly improve DiT-3D on ShapeNet. Experimental results on the ShapeNet dataset demonstrate that the proposed DiT-3D achieves state-of-the-art performance in high-fidelity and diverse 3D point cloud generation. In particular, our DiT-3D decreases the 1-Nearest Neighbor Accuracy of the state-of-the-art method by 4.59 and increases the Coverage metric by 3.51 when evaluated on Chamfer Distance.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d7557adc-04b5-4c0f-a7d2-e4c8743e0721Cited by top-tier papers57
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang et al.NeurIPS 2024 · 251 citations
- Learning-to-Cache: Accelerating Diffusion Transformer via Layer CachingXinyin Ma, Gongfan Fang, Michael Bi Mi, Xinchao WangNeurIPS 2024 · 167 citations
- Large-Vocabulary 3D Diffusion Model with TransformerZiang Cao, Fangzhou Hong, Tong Wu, Liang Pan et al.ICLR 2024 · 54 citations
- Wan-Move: Motion-controllable Video Generation via Latent Trajectory GuidanceRuihang Chu, Yefei He, Zhekai Chen, Shiwei Zhang et al.NeurIPS 2025 · 50 citations
- On Statistical Rates and Provably Efficient Criteria of Latent Diffusion Transformers (DiTs)Jerry Yao-Chieh Hu, Weimin Wu, Zhuoru Li, Sophia Pi et al.NeurIPS 2024 · 49 citations
Builds on22
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
Related papers
- Scaling Diffusion Mamba with Bidirectional SSMs for Efficient 3D Shape GenerationShentong MoAAAI 2025 · 4 citations
- Transformer-based Point Cloud Generation NetworkRui Xu, Le Hui, Yuehui Han, Jianjun Qian et al.ACM MM 2023 · 4 citations
- TIGER: Time-Varying Denoising Model for 3D Point Cloud Generation with Diffusion ProcessZhiyuan Ren, Minchul Kim, Feng Liu, Xiaoming LiuCVPR 2024 · 9 citations
- An End-to-End Transformer Model for 3D Object DetectionIshan Misra, Rohit Girdhar, Armand JoulinICCV 2021 · 602 citations
- Deep Point Cloud ReconstructionJaesung Choe, Byeongin Joung, François Rameau, Jaesik Park et al.ICLR 2022 · 28 citations
