PanoDiT: Panoramic Videos Generation with Diffusion Transformer
Muyang Zhang, Yuzhi Chen, Rongtao Xu, Changwei Wang, Jinming Yang, Weiliang Meng, Jianwei Guo, Huihuang Zhao, Xiaopeng Zhang
Abstract
As immersive experiences become increasingly popular, panoramic video has garnered significant attention in both research and applications. The high cost associated with capturing panoramic video underscores the need for efficient prompt-based generation methods. Although recent text-tovideo (T2V) diffusion techniques have shown potential in standard video generation, they face challenges when applied to panoramic videos due to substantial differences in content and motion patterns. In this paper, we propose Pan-oDiT, a framework that utilizes the Diffusion Transformer (DiT) architecture to generate panoramic videos from text descriptions. Unlike traditional methods that rely on UNetbased denoising, our method leverages a transformer architecture for denoising, incorporating both temporal and global attention mechanisms. This ensures coherent frame generation and smooth motion transitions, offering distinct advantages in long-horizon generation tasks. To further enhance motion and consistency in the generated videos, we introduce DTM-LoRA and two panoramic-specific losses. Compared to previous methods, our PanoDiT achieves state-of-the-art performance across various evaluation metrics and user study, with code is available in the supplementary material.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03d34dfd-2a61-4312-8d3b-a82c2d199ddaCited by top-tier papers3
- PanoWan: Lifting Diffusion Video Generation Models to 360° with Latitude/Longitude-aware MechanismsYifei Xia, Shuchen Weng, Siqi Yang, Jingqi Liu et al.NeurIPS 2025 · 24 citations
- CubeComposer: Spatio-Temporal Autoregressive 4K 360deg Video Generation from Perspective VideoLingen Li, Guangzhi Wang, Xiaoyu Li, Zhaoyang Zhang et al.CVPR 2026 · 10 citations
- PanFlow: Decoupled Motion Control for Panoramic Video GenerationCheng Zhang, Hanwen Liang, Donny Y. Chen, Qianyi Wu et al.AAAI 2026 · 1 citation
Builds on19
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and GenerationJunnan Li, Dongxu Li, Caiming Xiong, Steven C. H. HoiICML 2022 · 6,549 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
Related papers
- DiffPano: Scalable and Consistent Text to Panorama Generation with Spherical Epipolar-Aware DiffusionWeicai Ye, Chenhao Ji, Zheng Chen, Junyao Gao et al.NeurIPS 2024 · 45 citations
- 360DVD: Controllable Panorama Video Generation with 360-Degree Video Diffusion ModelQian Wang, Weiqi Li, Chong Mou, Xinhua Cheng et al.CVPR 2024 · 23 citations
- Generative Pre-trained Autoregressive Diffusion TransformerYuan Zhang, Jiacheng Jiang, Guoqing Ma, Zhiying Lu et al.NeurIPS 2025 · 19 citations
- DiT360: High-Fidelity Panoramic Image Generation via Hybrid TrainingHaoran Feng, Dizhe Zhang, Xiangtai Li, Bo Du et al.CVPR 2026 · 27 citations
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin et al.ICCV 2025
