Long Video Diffusion Generation with Segmented Cross-Attention and Content-Rich Video Data Curation
Xin Yan, Yuxuan Cai, Qiuyue Wang, Yuan Zhou, Wenhao Huang, Huan Yang
2025年份
7顶会引用
摘要
A young boy rides a bicycle down the long corridors of a towering ancient castle. The camera follows closely as he moves. As he exits the castle, the scene opens up to reveal a lush, vibrant garden filled with greenery and flowers, sunlight pouring over the landscape. An extreme close-up shot of an ant emerging from its nest. The camera pulls back revealing a neighborhood beyond the hill. Figure 1. Presto can generate long videos with rich content and long-range coherence.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- NarrLV: Towards a Comprehensive Narrative-Centric Evaluation for Long Video GenerationXiaokun Feng, Haiming Yu, Meiqi Wu, Shiyu Hu 等ICLR 2026 · 被引用 13 次
- Compositional Diffusion with Guided search for Long-Horizon PlanningUtkarsh A. Mishra, David He, Yongxin Chen, Danfei XuICLR 2026 · 被引用 9 次
- PhysVid: Physics Aware Local Conditioning for Generative Video ModelsSaurabh Pathak, Elahe Arani, Mykola Pechenizkiy, Bahram ZonoozCVPR 2026 · 被引用 6 次
- Narrative Weaver: Towards Controllable Long-Range Visual Consistency with Multi-Modal ConditioningZhengjian Yao, Yongzhi Li, Xinyuan Gao, Quan Chen 等CVPR 2026 · 被引用 3 次
- SteinsGate: Adding Causality to Diffusions for Long Video Generation via Path IntegralYufei Huang, Liangyu Yuan, Changxi Chi, Yunfan Liu 等ICLR 2026
它引用的顶会 Paper19
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu 等ICCV 2021 · 被引用 31,683 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion ModelsAlexander Quinn Nichol, Prafulla Dhariwal, Aditya Ramesh, Pranav Shyam 等ICML 2022 · 被引用 4,691 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- MPNet: Masked and Permuted Pre-training for Language UnderstandingKaitao Song, Xu Tan, Tao Qin, Jianfeng Lu 等NeurIPS 2020 · 被引用 1,957 次
相关 Paper
- TransPixeler: Advancing Text-to-Video Generation with TransparencyLuozhou Wang, Yijun Li, Zhifei Chen, Jui-Hsien Wang 等CVPR 2025
- Endless loops: detecting and animating periodic patterns in still imagesTavi Halperin, Hanit Hakim, Orestis Vantzos, Gershon Hochman 等SIGGRAPH 2021 · 被引用 14 次
- VideoCrafter2: Overcoming Data Limitations for High-Quality Video Diffusion ModelsHaoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia 等CVPR 2024
- Dream3D: Zero-Shot Text-to-3D Synthesis Using 3D Shape Prior and Text-to-Image Diffusion ModelsJiale Xu, Xintao Wang, Weihao Cheng, Yan-Pei Cao 等CVPR 2023
- Tora: Trajectory-oriented Diffusion Transformer for Video GenerationZhenghao Zhang, Junchao Liao, Menghao Li, Zuozhuo Dai 等CVPR 2025
