Controlling Space and Time with Diffusion Models
Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasacchi, David J. Fleet
Abstract
We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works which generally operate in limited domains (e.g., object centric). 4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned videoto-video translation, which we illustrate qualitatively on a variety of scenes. See https://4d-diffusion.github.io for video samples. Input 360 • rotation * Equal contribution. Unlike the typical setting of Neural Radiance Fields (Mildenhall et al., 2021) (NeRF) where tens-tohundreds of images are used as input for 3D reconstruction, pose-conditional diffusion models for NVS aim to extrapolate plausible, diverse, 3D consistent samples with as few as a single image input. Conditioning diffusion models on an image and relative camera pose was introduced by Watson et al. * 4DiM does not require sequential temporal ordering as in video models as the architecture is permutationequivariant over frames. All N images (conditioning and generated) are processed by the diffusion model. * For PNVS, we follow Yu et al. (2023a) and use a Markovian sliding window for sampling, as they find it is the stronger than stochastic conditioning (Watson et al., 2022).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 26241a4f-82ae-4127-9d1d-7cbb54de61d3Cited by top-tier papers37
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory ControlXiao Fu, Xintao Wang, Xian Liu, Jianhong Bai et al.ICLR 2026 · 37 citations
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang et al.ICLR 2026 · 33 citations
- Stable Virtual Camera: Generative View Synthesis with Diffusion ModelsJensen Zhou, Hang Gao, Vikram Voleti, Aaryaman Vasishta et al.ICCV 2025 · 25 citations
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video GuidanceZun Wang, Jaemin Cho, Jialu Li, Han Lin et al.ICML 2026 · 19 citations
- Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular VideosKaihua Chen, Tarasha Khurana, Deva RamananNeurIPS 2025 · 17 citations
Builds on29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoYanqin Jiang, Li Zhang, Jin Gao, Weiming Hu et al.ICLR 2024 · 120 citations
- BulletTime: Decoupled Control of Time and Camera Pose for Video GenerationYiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu et al.CVPR 2026 · 15 citations
- Novel View Synthesis with Diffusion ModelsDaniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho et al.ICLR 2023 · 63 citations
- SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View ConsistencyYiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang et al.ICLR 2025
- Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video DiffusionWenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen et al.ICCV 2025 · 11 citations
