ViewPoint: Panoramic Video Generation with Pretrained Diffusion Models
Zixun Fang, Kai Zhu, Zhiheng Liu, Yu Liu, Wei Zhai, Yang Cao, Zheng-Jun Zha
Abstract
Panoramic video generation aims to synthesize 360-degree immersive videos, holding significant importance in the fields of VR, world models, and spatial intelligence. Existing works fail to synthesize high-quality panoramic videos due to the inherent modality gap between panoramic data and perspective data, which constitutes the majority of the training data for modern diffusion models. In this paper, we propose a novel framework utilizing pretrained perspective video models for generating panoramic videos. Specifically, we design a novel panorama representation named ViewPoint map, which possesses global spatial continuity and fine-grained visual details simultaneously. With our proposed Pano-Perspective attention mechanism, the model benefits from pretrained perspective priors and captures the panoramic spatial correlations of the ViewPoint map effectively. Extensive experiments demonstrate that our method can synthesize highly dynamic and spatially consistent panoramic videos, achieving state-of-the-art performance and surpassing previous methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- WorldWeaver: Generating Long-Horizon Video Worlds via Rich PerceptionZhiheng Liu, Xueqing Deng, Shoufa Chen, Angtian Wang et al.NeurIPS 2025 · 15 citations
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective VideoLingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang et al.CVPR 2026 · 10 citations
- Diffusion Guided Chain-of-Vision for Large Autoregressive Vision ModelsXinyang Wang, Kecheng Zheng, Minfeng Zhu, Wei Wu et al.CVPR 2026
- When Numbers Speak: Aligning Textual Numerals and Visual Instances in Text-to-Video Diffusion ModelsZhengyang Sun, Yu Chen, Xin Zhou, Xiaofan Li et al.CVPR 2026
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 11,743 citations
- Directly Denoising Diffusion ModelsDan Zhang, Jingjing Wang, Feng LuoICML 2024 · 11,724 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
Related papers
- Look Beyond: Two-Stage Scene View Generation via Panorama and Video DiffusionXueyang Kang, Zhengkang Xiang, Zezheng Zhang, Kourosh KhoshelhamACM MM 2025
- OmniRoam: World Wandering via Long-Horizon Panoramic Video GenerationYuheng Liu, Xin Lin, Xinke Li, Baihan Yang et al.SIGGRAPH 2026 · 1 citation
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware MechanismsYifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu et al.NeurIPS 2025 · 24 citations
- 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelQian Wang, Weiqi Li, Chong Mou, Xinhua Cheng et al.CVPR 2024 · 23 citations
- HoloTime: Taming Video Diffusion Models for Panoramic 4D Scene GenerationHaiyang Zhou, Wangbo Yu, Jiawen Guan, Xinhua Cheng et al.ACM MM 2025 · 4 citations
