Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion Transformers
Minghao Yin, Wenbo Hu, Jiale Xu, Ying Shan, Kai Han
摘要
Recent breakthroughs in 3D generative modeling have yielded remarkable progress in static shape synthesis, yet high-fidelity dynamic 4D generation remains elusive, hindered by temporal artifacts and prohibitive computational demand. We present Sculpt4D, a native 4D generative framework that seamlessly integrates efficient temporal modeling into a pretrained 3D Diffusion Transformer (Hunyuan3D 2.1), thereby mitigating the scarcity of 4D training data. At its core lies a Block Sparse Attention mechanism that preserves object identity by anchoring to the initial frame while capturing rich motion dynamics via a time-decaying sparse mask. This design faithfully models complex spatiotemporal dependencies with high fidelity, while sidestepping the quadratic overhead of full attention and reducing network total computation by 56%. Consequently, Sculpt4D establishes a new state-of-the-art in temporally coherent 4D synthesis and charts a path toward efficient and scalable 4D generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper41
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Big Bird: Transformers for Longer SequencesManzil Zaheer, Guru Guruganesh, Kumar Avinava Dubey, Joshua Ainslie 等NeurIPS 2020 · 被引用 3,159 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- Efficient Geometry-aware 3D Generative Adversarial NetworksEric R. Chan, Connor Z. Lin, Matthew A. Chan, Koki Nagano 等CVPR 2022 · 被引用 984 次
- GET3D: A Generative Model of High Quality 3D Textured Shapes Learned from ImagesJun Gao, Tianchang Shen, Zian Wang, Wenzheng Chen 等NeurIPS 2022 · 被引用 661 次
相关 Paper
- Turbo4DGen: Ultra-Fast Acceleration for 4D GenerationYuanbin Man, Ying Huang, Zhile Ren, Miao YinICML 2026
- ShapeGen4D: Towards High Quality 4D Shape Generation from VideosJiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen, Jiaxu Zou 等ICLR 2026 · 被引用 21 次
- Interspatial Attention for Efficient 4D Human Video GenerationRuizhi Shao, Yinghao Xu, Yujun Shen, Ceyuan Yang 等SIGGRAPH 2025 · 被引用 2 次
- Direct3D-S2: Gigascale 3D Generation Made Easy with Spatial Sparse AttentionShuang Wu, Youtian Lin, Feihu Zhang, Yifei Zeng 等NeurIPS 2025 · 被引用 114 次
- DFSAttn: Dynamic Fine-grained Sparse Attention for Efficient Video GenerationJie Hu, Zixiang Gao, Yutong He, Kun YuanICML 2026
