PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware Mechanisms
Yifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu, Chengxuan Zhu, Minggui Teng, Zijian Jia, Han Jiang, Boxin Shi
摘要
Panoramic video generation enables immersive 360 content creation, valuable in applications that demand scene-consistent world exploration. However, existing panoramic video generation models struggle to leverage pre-trained generative priors from conventional text-to-video models for high-quality and diverse panoramic videos generation, due to limited dataset scale and the gap in spatial feature representations. In this paper, we introduce PanoWan to effectively lift pre-trained text-to-video models to the panoramic domain, equipped with minimal modules. PanoWan employs latitude-aware sampling to avoid latitudinal distortion, while its rotated semantic denoising and padded pixel-wise decoding ensure seamless transitions at longitude boundaries. To provide sufficient panoramic videos for learning these lifted representations, we contribute PanoVid, a high-quality panoramic video dataset with captions and diverse scenarios. Consequently, PanoWan achieves state-of-the-art performance in panoramic video generation and demonstrates robustness for zero-shot downstream tasks. Our project page is available at https://panowan.variantconst.com.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Unified Camera Positional Encoding for Controlled Video GenerationCheng Zhang, Boying Li, Meng Wei, Yan-Pei Cao 等CVPR 2026 · 被引用 38 次
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective VideoLingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang 等CVPR 2026 · 被引用 10 次
它引用的顶会 Paper17
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Is Space-Time Attention All You Need for Video Understanding?Gedas Bertasius, Heng Wang, Lorenzo TorresaniICML 2021 · 被引用 2,927 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
- Q-Align: Teaching LMMs for Visual Scoring via Discrete Text-Defined LevelsHaoning Wu, Zicheng Zhang, Weixia Zhang, Chaofeng Chen 等ICML 2024 · 被引用 499 次
相关 Paper
- 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelQian Wang, Weiqi Li, Chong Mou, Xinhua Cheng 等CVPR 2024 · 被引用 23 次
- OmniRoam: World Wandering via Long-Horizon Panoramic Video GenerationYuheng Liu, Xin Lin, Xinke Li, Baihan Yang 等SIGGRAPH 2026 · 被引用 1 次
- ViewPoint: Panoramic Video Generation with Pretrained Diffusion ModelsZixun Fang, Kai Zhu, Zhiheng Liu, Yu Liu 等NeurIPS 2025 · 被引用 2 次
- DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionWeicai Ye, Chenhao Ji, Zheng Chen, Junyao Gao 等NeurIPS 2024 · 被引用 45 次
- PanoDiT: Panoramic Videos Generation with Diffusion TransformerMuyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang 等AAAI 2025 · 被引用 6 次
