PanoDiT: Panoramic Videos Generation with Diffusion Transformer
Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang, Jinming Yang, Weiliang Meng, Jianwei Guo, Huihuang Zhao, Xiaopeng Zhang
摘要
As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-tovideo (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose Pan-oDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNetbased denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware MechanismsYifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu 等NeurIPS 2025 · 被引用 24 次
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective VideoLingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang 等CVPR 2026 · 被引用 10 次
- PanFlow: Decoupled Motion Control for Panoramic Video GenerationCheng Zhang, Hanwen Liang, Donny Y. Chen, Qianyi Wu 等AAAI 2026 · 被引用 1 次
它引用的顶会 Paper19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 被引用 6,549 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
相关 Paper
- DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionWeicai Ye, Chenhao Ji, Zheng Chen, Junyao Gao 等NeurIPS 2024 · 被引用 45 次
- 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelQian Wang, Weiqi Li, Chong Mou, Xinhua Cheng 等CVPR 2024 · 被引用 23 次
- Generative Pre-trained Autoregressive Diffusion TransformerYuan Zhang, Jiacheng Jiang, Guoqing Ma, Zhiying Lu 等NeurIPS 2025 · 被引用 19 次
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid TrainingHaoran Feng, Dizhe Zhang, Xiangtai Li, Bo Du 等CVPR 2026 · 被引用 27 次
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin 等ICCV 2025
