WorldReel: 4D Video Generation with Consistent Geometry and Motion Modeling
Shaoheng Fang, Hanwen Jiang, Yunpeng Bai, Niloy J. Mitra, Qixing Huang
Abstract
Recent video generators achieve striking photorealism, yet remain fundamentally inconsistent in 3D. We present WorldReel, a 4D video generator that is natively spatio-temporally consistent. WorldReel jointly produces RGB frames together with 4D scene representations, including pointmaps, camera trajectory, and dense flow mapping, enabling coherent geometry and appearance modeling over time. Our explicit 4D representation enforces a single underlying scene that persists across viewpoints and dynamic content, yielding videos that remain consistent even under large non-rigid motion and significant camera movement.We train WorldReel by carefully combining synthetic and real data: synthetic data provides precise 4D supervision (geometry, motion, and camera), while real videos contribute visual diversity and realism. This blend allows WorldReel to generalize to in-the-wild footage while preserving strong geometric fidelity.Extensive experiments demonstrate that WorldReel sets a new state-of-the-art for consistent video generation with dynamic scenes and moving cameras, improving metrics of geometric consistency, motion coherence, and reducing view-time artifacts over competing methods. We believe that WorldReel brings video generation closer to 4D-consistent world modeling, where agents can render, interact, and reason about scenes through a single and stable spatiotemporal representation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b7681480-58f3-4857-99ac-dae0d3a0e8b0Builds on51
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Vision Transformers for Dense PredictionRené Ranftl, Alexey Bochkovskiy, Vladlen KoltunICCV 2021 · 2,647 citations
Related papers
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng et al.ICLR 2026 · 28 citations
- Restage4D: Reanimating Deformable 3D Reconstruction from a Single VideoJixuan He, Chieh Hubert Lin, Lu Qi, Ming-Hsuan YangNeurIPS 2025 · 2 citations
- ShapeGen4D: Towards High Quality 4D Shape Generation from VideosJiraphon Yenphraphai, Ashkan Mirzaei, Jianqi Chen, Jiaxu Zou et al.ICLR 2026 · 21 citations
- VerseCrafter: Dynamic Realistic Video World Model with 4D Geometric ControlSixiao Zheng, Minghao Yin, Wenbo Hu, Xiaoyu Li et al.CVPR 2026 · 27 citations
- Generating Long Videos of Dynamic ScenesTim Brooks, Janne Hellsten, Miika Aittala, Ting-Chun Wang et al.NeurIPS 2022 · 152 citations
