Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
Xin Yan, Yuxuan Cai, Qiuyue Wang, Yuan Zhou, Wenhao Huang, Huan Yang
2025Year
7Top-tier citations
Abstract
A young boy rides a bicycle down the long corridors of a towering ancient castle. The camera follows closely as he moves. As he exits the castle, the scene opens up to reveal a lush, vibrant garden filled with greenery and flowers, sunlight pouring over the landscape. An extreme close-up shot of an ant emerging from its nest. The camera pulls back revealing a neighborhood beyond the hill. Figure 1. Presto can generate long videos with rich content and long-range coherence.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video GenerationXiaokun Feng, Haiming Yu, Meiqi Wu, Shiyu Hu et al.ICLR 2026 · 13 citations
- Compositional Diffusion with Guided search for Long-Horizon PlanningUtkarsh A. Mishra, David He, Yongxin Chen, Danfei XuICLR 2026 · 9 citations
- PhysVid: Physics Aware Local Conditioning for Generative Video ModelsSaurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram ZonoozCVPR 2026 · 6 citations
- Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal ConditioningZhengjian Yao, Yongzhi Li, Xinyuan Gao, Quan Chen et al.CVPR 2026 · 3 citations
- SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path IntegralYufei Huang, Liangyu Yuan, Changxi Chi, Yunfan Liu et al.ICLR 2026
Builds on19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam et al.ICML 2022 · 4,691 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu et al.NeurIPS 2020 · 1,957 citations
Related papers
- TransPixeler: Advancing Text-to-Video Generation with TransparencyLuozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang et al.CVPR 2025
- Endless loops: detecting and animating periodic patterns in still imagesTavi Halperin, Hanit Hakim, Orestis Vantzos, Gershon Hochman et al.SIGGRAPH 2021 · 14 citations
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion ModelsHaoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia et al.CVPR 2024
- Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion ModelsJiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao et al.CVPR 2023
- Tora: Trajectory-oriented Diffusion Transformer for Video GenerationZhenghao Zhang, Junchao Liao, Menghao Li, Zuozhuo Dai et al.CVPR 2025
