SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View Consistency
Yiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang, Varun Jampani
摘要
We present Stable Video 4D (SV4D) -a latent video diffusion model for multiframe and multi-view consistent dynamic 3D content generation. Unlike previous methods that rely on separately trained generative models for video generation and novel view synthesis, we design a unified diffusion model to generate novel view videos of dynamic 3D objects. Specifically, given a monocular reference video, SV4D generates novel views for each video frame that are temporally consistent. We then use the generated novel view videos to optimize an implicit 4D representation (dynamic NeRF) efficiently, without the need for cumbersome SDS-based optimization used in most prior works. To train our unified novel view video generation model, we curate a dynamic 3D object dataset from the existing Objaverse dataset. Extensive experimental results on multiple datasets and user studies demonstrate SV4D's state-of-the-art performance on novel-view video synthesis as well as 4D generation compared to prior works. Project page: https://sv4d.github.io .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper68
- 4DGT: Learning a 4D Gaussian Transformer Using Real-World Monocular VideosZhen Xu, Zhengqin Li, Zhao Dong, Xiaowei Zhou 等NeurIPS 2025 · 被引用 51 次
- Puppeteer: Rig and Animate Your 3D ModelsChaoyue Song, Xiu Li, Fan Yang, Zhongcong Xu 等NeurIPS 2025 · 被引用 48 次
- World-In-World: World Models in a Closed-Loop WorldJiahan Zhang, Muqing Jiang, Nanru Dai, Taiming Lu 等ICLR 2026 · 被引用 46 次
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang 等ICLR 2026 · 被引用 33 次
- Recammaster: Camera-Controlled Generative Rendering From a Single VideoJianhong Bai, Menghan Xia, Xiao Fu, Xintao Wang 等ICCV 2025 · 被引用 33 次
它引用的顶会 Paper59
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 被引用 5,687 次
- Instant neural graphics primitives with a multiresolution hash encodingThomas Müller, Alex Evans, Christoph Schied, Alexander KellerSIGGRAPH 2022 · 被引用 4,089 次
相关 Paper
- 4Diffusion: Multi-view Video Diffusion Model for 4D GenerationHaiyu Zhang, Xinyuan Chen, Yaohui Wang, Xihui Liu 等NeurIPS 2024 · 被引用 119 次
- SV4D 2.0: Enhancing Spatio-Temporal Consistency in Multi-View Video Diffusion for High-Quality 4D GenerationChun-Han Yao, Yiming Xie, Vikram Voleti, Huaizu Jiang 等ICCV 2025 · 被引用 5 次
- Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoYanqin Jiang, Li Zhang, Jin Gao, Weiming Hu 等ICLR 2024 · 被引用 120 次
- Text-To-4D Dynamic Scene GenerationUriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual 等ICML 2023 · 被引用 234 次
- Dynamic View Synthesis from Dynamic Monocular VideoChen Gao, Ayush Saraf, Johannes Kopf, Jia-Bin HuangICCV 2021 · 被引用 522 次
