HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene Generation
Haiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng, Yonghong Tian, Li Yuan
Abstract
The rapid advancement of diffusion models holds the promise of revolutionizing the application of VR and AR technologies, which typically require scene-level 4D assets for user experience. Nonetheless, existing diffusion models predominantly concentrate on modeling static 3D scenes or object-level dynamics, constraining their capacity to provide truly immersive experiences. To address this issue, we propose HoloTime, a framework that integrates video diffusion models to generate panoramic videos from a single prompt or reference image, along with a 360-degree 4D scene reconstruction method that seamlessly transforms the generated panoramic video into 4D assets, enabling a fully immersive 4D experience for users. Specifically, to tame video diffusion models for generating high-fidelity panoramic videos, we introduce the 360World dataset, the first comprehensive collection of panoramic videos suitable for downstream 4D scene reconstruction tasks. With this curated dataset, we propose Panoramic Animator, a two-stage image-to-video diffusion model that can convert panoramic images into high-quality panoramic videos. Following this, we present Panoramic Space-Time Reconstruction, which leverages a space-time depth estimation method to transform the generated panoramic videos into 4D point clouds, enabling the optimization of a holistic 4D Gaussian Splatting representation to reconstruct spatially and temporally consistent 4D scenes. To validate the efficacy of our method, we conducted a comparative analysis with existing approaches, revealing its superiority in both panoramic video generation and 4D scene reconstruction. This demonstrates our method's capability to create more engaging and realistic immersive environments, thereby enhancing user experiences in VR and AR applications.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 314aaf91-9b29-4a8a-a815-87b970f5ed96Cited by top-tier papers6
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware MechanismsYifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu et al.NeurIPS 2025 · 24 citations
- MATRIX: Mask Track Alignment for Interaction-aware Video GenerationSiyoon Jin, Seongchan Kim, Jae Ho Lee, Dahyun Chung et al.ICLR 2026 · 4 citations
- EasyVFX: Frequency-Driven Decoupling for Resource-Efficient VFX GenerationYue Ma, Xu Ye, Qinghe Wang, Yucheng Wang et al.SIGGRAPH 2026 · 2 citations
- PanFlow: Decoupled Motion Control for Panoramic Video GenerationCheng Zhang, Hanwen Liang, Donny Y. Chen, Qianyi Wu et al.AAAI 2026 · 1 citation
- EA3D: Event-Augmented 3D Diffusion for Generalizable Novel View SynthesisWangbo Yu, Chaoran Feng, Jianing Li, Aofan Zhang et al.ICLR 2026
Builds on43
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- 3D Gaussian Splatting for Real-Time Radiance Field RenderingBernhard Kerbl, Georgios Kopanas, Thomas Leimkühler, George DrettakisSIGGRAPH 2023 · 5,687 citations
Related papers
- 4K4DGen: Panoramic 4D Generation at 4K ResolutionRenjie Li, Panwang Pan, Bangbang Yang, Dejia Xu et al.ICLR 2025
- ViewPoint: Panoramic Video Generation with Pretrained Diffusion ModelsZixun Fang, Kai Zhu, Zhiheng Liu, Yu Liu et al.NeurIPS 2025 · 2 citations
- TiP4GEN: Text to Immersive Panorama 4D Scene GenerationKe Xing, Hanwen Liang, Dejia Xu, Yuyang Yin et al.ACM MM 2025 · 2 citations
- 360Explorer: Exploring 4D Controllable World in Panoramic VideosXinhua Cheng, Haiyang Zhou, Wangbo Yu, Tanghui Jia et al.AAAI 2026
- 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelQian Wang, Weiqi Li, Chong Mou, Xinhua Cheng et al.CVPR 2024 · 23 citations
