Turbo4DGen: Ultra-Fast Acceleration for 4D Generation
Yuanbin Man, Ying Huang, Zhile Ren, Miao Yin
摘要
4D generation, or dynamic 3D content generation, integrates spatial, temporal, and view dimensions to model realistic dynamic scenes, playing a foundational role in advancing world models and physical AI. However, maintaining long-chain consistency across both frames and viewpoints through the unique spatio-camera-motion (SCM) attention mechanism introduces substantial computational and memory overhead, often leading to out-of-memory (OOM) failures and prohibitive generation times. To address these challenges, we propose Turbo4DGen, an ultra-fast acceleration framework for diffusion-based multi-view 4D content generation. Turbo4DGen introduces a spatiotemporal cache mechanism that persistently reuses intermediate attention across denoising steps, combined with dynamically semantic-aware attention pruning and an adaptive SCM chain bypass scheduler, to drastically reduce redundant SCM attention computation. Our experimental results show that Turbo4DGen achieves an average 9.7 speedup without quality degradation on the ObjaverseDy and Consistent4D datasets. To the best of our knowledge, Turbo4DGen is the first dedicated acceleration framework for 4D generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- FlexGen: High-Throughput Generative Inference of Large Language Models with a Single GPUYing Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li 等ICML 2023 · 被引用 683 次
- DreamFusion: Text-to-3D using 2D DiffusionBen Poole, Ajay Jain, Jonathan T. Barron, Ben MildenhallICLR 2023 · 被引用 463 次
相关 Paper
- Diffusion2: Dynamic 3D Content Generation via Score Composition of Video and Multi-view Diffusion ModelsZeyu Yang, Zijie Pan, Chun Gu, Li ZhangICLR 2025
- RAPID: Reusing Attention Sparsity with Inter-step Adaptation for Efficient Video DiffusionShangran Lin, Lu Lu, Jian Chen, Qiang LiuCVPR 2026
- Sculpt4D: Generating 4D Shapes via Sparse-Attention Diffusion TransformersMinghao Yin, Wenbo Hu, Jiale Xu, Ying Shan 等CVPR 2026 · 被引用 3 次
- Fused View-Time Attention and Feedforward Reconstruction for 4D Scene GenerationChaoyang Wang, Ashkan Mirzaei, Vidit Goel, Willi Menapace 等NeurIPS 2025 · 被引用 12 次
- Diffusion4D: Fast Spatial-temporal Consistent 4D generation via Video Diffusion ModelsHanwen Liang, Yuyang Yin, Dejia Xu, Hanxue Liang 等NeurIPS 2024 · 被引用 116 次
