4D-fy: Text-to-4D Generation Using Hybrid Score Distillation Sampling
Sherwin Bahmani, Ivan Skorokhodov, Victor Rong, Gordon Wetzstein, Leonidas J. Guibas, Peter Wonka, Sergey Tulyakov, Jeong Joon Park, Andrea Tagliasacchi, David B. Lindell
2024Year
3Top-tier citations
Abstract
7 SFU 8 Google "a space shuttle launching" "a crocodile playing a drum set" time viewpoint Figure 1. Text-to-4D Synthesis. We present 4D-fy, a technique that synthesizes 4D (i.e., dynamic 3D) scenes from a text prompt. We show scenes generated from two text prompts for different viewpoints (vertical dimension) at different time steps (horizontal dimension). Video results can be viewed on our website: https://sherwinbahmani.github.io/4dfy .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control ConditionsYuanhao Cai, He Zhang, Xi Chen, Jinbo Xing et al.NeurIPS 2025 · 19 citations
- Restage4D: Reanimating Deformable 3D Reconstruction from a Single VideoJixuan He, Chieh Hubert Lin, Lu Qi, Ming-Hsuan YangNeurIPS 2025 · 2 citations
- Wonderland: Navigating 3D Scenes from a Single ImageHanwen Liang, Junli Cao, Vidit Goel, Guocheng Qian et al.CVPR 2025
Builds on39
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
Related papers
- CAT4D: Create Anything in 4D with Multi-View Video Diffusion ModelsRundi Wu, Ruiqi Gao, Ben Poole, Alex Trevithick et al.CVPR 2025
- Text-To-4D Dynamic Scene GenerationUriel Singer, Shelly Sheynin, Adam Polyak, Oron Ashual et al.ICML 2023 · 234 citations
- S-DyRF: Reference-Based Stylized Radiance Fields for Dynamic ScenesXingyi Li, Zhiguo Cao, Yizheng Wu, Kewei Wang et al.CVPR 2024
- A Unified Approach for Text-and Image-Guided 4D Scene GenerationYufeng Zheng, Xueting Li, Koki Nagano, Sifei Liu et al.CVPR 2024 · 20 citations
- 4Real: Towards Photorealistic 4D Scene Generation via Video Diffusion ModelsHeng Yu, Chaoyang Wang, Peiye Zhuang, Willi Menapace et al.NeurIPS 2024 · 76 citations
