Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse Attention
Shuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng, Yikang Yang, Yajie Bao, Jiachen Qian, Siyu Zhu, Xun Cao, Philip H. S. Torr, Yao Yao
摘要
Generating high-resolution 3D shapes using volumetric representations such as Signed Distance Functions (SDFs) presents substantial computational and memory challenges. We introduce Direct3D-S2, a scalable 3D generation framework based on sparse volumes that achieves superior output quality with dramatically reduced training costs. Our key innovation is the Spatial Sparse Attention (SSA) mechanism, which greatly enhances the efficiency of Diffusion Transformer (DiT) computations on sparse volumetric data. SSA allows the model to effectively process large token sets within sparse volumes, substantially reducing computational overhead and achieving a 3.9x speedup in the forward pass and a 9.6x speedup in the backward pass. Our framework also includes a variational autoencoder (VAE) that maintains a consistent sparse volumetric format across input, latent, and output stages. Compared to previous methods with heterogeneous representations in 3D VAE, this unified design significantly improves training efficiency and stability. Our model is trained on public available datasets, and experiments demonstrate that Direct3D-S2 not only surpasses state-of-the-art methods in generation quality and efficiency, but also enables training at 1024 resolution using only 8 GPUs, a task typically requiring at least 32 GPUs for volumetric representations at 256 resolution, thus making gigascale 3D generation both practical and accessible. Project page: https://www.neural4d.com/research/direct3d-s2.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper36
- SAM 3D: 3Dfy Anything in ImagesXingyu Chen, Fu-Jen Chu, Pierre Gleize, Kevin J Liang 等CVPR 2026 · 被引用 280 次
- Native and Compact Structured Latents for 3D GenerationJianfeng Xiang, Xiaoxue Chen, Sicheng Xu, Ruicheng Wang 等CVPR 2026 · 被引用 177 次
- LATTICE: Democratize High-Fidelity 3D Generation at ScaleZeqiang Lai, Yunfei Zhao, Zibo Zhao, Haolin Liu 等CVPR 2026 · 被引用 46 次
- ShapeR: Robust Conditional 3D Shape Generation from Casual CapturesYawar Siddiqui, Duncan P. Frost, Samir Aroudj, Armen Avetisyan 等CVPR 2026 · 被引用 19 次
- FullPart: Generating each 3D Part at Full ResolutionLihe Ding, Shaocong Dong, Yaokun Li, Chenjian Gao 等ICLR 2026 · 被引用 17 次
它引用的顶会 Paper26
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Transformers are RNNs: Fast Autoregressive Transformers with Linear AttentionAngelos Katharopoulos, Apoorv Vyas, Nikolaos Pappas, François FleuretICML 2020 · 被引用 2,665 次
- FlashAttention-2: Faster Attention with Better Parallelism and Work PartitioningTri DaoICLR 2024 · 被引用 2,600 次
相关 Paper
- DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video GenerationJie Hu, Zixiang Gao, Yutong He, Kun YuanICML 2026
- Sparse Video-Gen: Accelerating Video Diffusion Transformers with Spatial-Temporal SparsityHaocheng Xi, Shuo Yang, Yilong Zhao, Chenfeng Xu 等ICML 2025
- Direct3D: Scalable Image-to-3D Generation via 3D Latent Diffusion TransformerShuang Wu, Youtian Lin, Yifei Zeng, Feihu Zhang 等NeurIPS 2024 · 被引用 251 次
- Faster Video Diffusion with Trainable Sparse AttentionPeiyuan Zhang, Yongqi Chen, Haofeng Huang, Will Lin 等NeurIPS 2025 · 被引用 6 次
- Interspatial Attention for Efficient 4D Human Video GenerationRuizhi Shao, Yinghao Xu, Yujun Shen, Ceyuan Yang 等SIGGRAPH 2025 · 被引用 2 次
