Controlling Space and Time with Diffusion Models
Daniel Watson, Saurabh Saxena, Lala Li, Andrea Tagliasacchi, David J. Fleet
摘要
We present 4DiM, a cascaded diffusion model for 4D novel view synthesis (NVS), supporting generation with arbitrary camera trajectories and timestamps, in natural scenes, conditioned on one or more images. With a novel architecture and sampling procedure, we enable training on a mixture of 3D (with camera pose), 4D (pose+time) and video (time but no pose) data, which greatly improves generalization to unseen images and camera pose trajectories over prior works which generally operate in limited domains (e.g., object centric). 4DiM is the first-ever NVS method with intuitive metric-scale camera pose control enabled by our novel calibration pipeline for structure-from-motion-posed data. Experiments demonstrate that 4DiM outperforms prior 3D NVS models both in terms of image fidelity and pose alignment, while also enabling the generation of scene dynamics. 4DiM provides a general framework for a variety of tasks including single-image-to-3D, two-image-to-video (interpolation and extrapolation), and pose-conditioned videoto-video translation, which we illustrate qualitatively on a variety of scenes. See https://4d-diffusion.github.io for video samples. Input 360 • rotation * Equal contribution. Unlike the typical setting of Neural Radiance Fields (Mildenhall et al., 2021) (NeRF) where tens-tohundreds of images are used as input for 3D reconstruction, pose-conditional diffusion models for NVS aim to extrapolate plausible, diverse, 3D consistent samples with as few as a single image input. Conditioning diffusion models on an image and relative camera pose was introduced by Watson et al. * 4DiM does not require sequential temporal ordering as in video models as the architecture is permutationequivariant over frames. All N images (conditioning and generated) are processed by the diffusion model. * For PNVS, we follow Yu et al. (2023a) and use a Markovian sliding window for sampling, as they find it is the stronger than stochastic conditioning (Watson et al., 2022).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper37
- Learning Video Generation for Robotic Manipulation with Collaborative Trajectory ControlXiao Fu, Xintao Wang, Xian Liu, Jianhong Bai 等ICLR 2026 · 被引用 37 次
- Lyra: Generative 3D Scene Reconstruction via Video Diffusion Model Self-DistillationSherwin Bahmani, Tianchang Shen, Jiawei Ren, Jiahui Huang 等ICLR 2026 · 被引用 33 次
- Stable Virtual Camera: Generative View Synthesis with Diffusion ModelsJensen Zhou, Hang Gao, Vikram Voleti, Aaryaman Vasishta 等ICCV 2025 · 被引用 25 次
- EPiC: Efficient Video Camera Control Learning with Precise Anchor-Video GuidanceZun Wang, Jaemin Cho, Jialu Li, Han Lin 等ICML 2026 · 被引用 19 次
- Reconstruct, Inpaint, Test-Time Finetune: Dynamic Novel-view Synthesis from Monocular VideosKaihua Chen, Tarasha Khurana, Deva RamananNeurIPS 2025 · 被引用 17 次
它引用的顶会 Paper29
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- Consistent4D: Consistent 360° Dynamic Object Generation from Monocular VideoYanqin Jiang, Li Zhang, Jin Gao, Weiming Hu 等ICLR 2024 · 被引用 120 次
- BulletTime: Decoupled Control of Time and Camera Pose for Video GenerationYiming Wang, Qihang Zhang, Shengqu Cai, Tong Wu 等CVPR 2026 · 被引用 15 次
- Novel View Synthesis with Diffusion ModelsDaniel Watson, William Chan, Ricardo Martin-Brualla, Jonathan Ho 等ICLR 2023 · 被引用 63 次
- SV4D: Dynamic 3D Content Generation with Multi-Frame and Multi-View ConsistencyYiming Xie, Chun-Han Yao, Vikram Voleti, Huaizu Jiang 等ICLR 2025
- Dimensionx: Create Any 3D and 4D Scenes From a Single Image With Decoupled Video DiffusionWenqiang Sun, Shuo Chen, Fangfu Liu, Zilong Chen 等ICCV 2025 · 被引用 11 次
