UltraViCo: Breaking Extrapolation Limits in Video Diffusion Transformers
Min Zhao, Hongzhou Zhu, Yingze Wang, Bokai Yan, Jintao Zhang, Guande He, Ling Yang, Chongxuan Li, Jun Zhu
Abstract
Despite advances, video diffusion transformers still struggle to generalize beyond their training length, a challenge we term video length extrapolation. We identify two failure modes: model-specific periodic content repetition and a universal quality degradation. Prior works attempt to solve repetition via positional encodings, overlooking quality degradation and achieving only limited extrapolation. In this paper, we revisit this challenge from a more fundamental view-attention maps, which directly govern how context influences outputs. We identify that both failure modes arise from a unified cause: attention dispersion, where tokens beyond the training window dilute learned attention patterns. This leads to quality degradation and repetition emerges as a special case when this dispersion becomes structured into periodic attention patterns, induced by harmonic properties of positional encodings. Building on this insight, we propose UltraViCo, a training-free, plug-and-play method that suppresses attention for tokens beyond the training window via a constant decay factor. By jointly addressing both failure modes, we outperform a broad set of baselines largely across models and extrapolation ratios, pushing the extrapolation limit from 2× to 4×. Remarkably, it improves Dynamic Degree and Imaging Quality by 233% and 40.5% over the previous best method at 4× extrapolation. Furthermore, our method generalizes seamlessly to downstream tasks such as controllable video synthesis and editing. Project page is available at https://thu-ml.github.io/UltraViCo.github.io/ .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 05850100-eb07-47ec-9b49-add2a5e197f2Cited by top-tier papers1
Ask how each one uses itBuilds on34
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-AwarenessTri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra et al.NeurIPS 2022 · 5,493 citations
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar et al.ICLR 2021 · 1,270 citations
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
Related papers
- RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion TransformersMin Zhao, Guande He, Yixiao Chen, Hongzhou Zhu et al.ICML 2025
- Free-Lunch Long Video Generation via Layer-Adaptive O.O.D CorrectionJiahao Tian, Chenxi Song, Wei Cheng, Chi ZhangCVPR 2026 · 3 citations
- FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal AttentionYu Lu, Yuanzhi Liang, Linchao Zhu, Yi YangNeurIPS 2024 · 101 citations
- LongDiff: Training-Free Long Video Generation in One GoZhuoling Li, Hossein Rahmani, Qiuhong Ke, Jun LiuCVPR 2025
- DyPE: Dynamic Position Extrapolation for Ultra High Resolution DiffusionNoam Issachar, Guy Yariv, Sagie Benaim, Yossi Adi et al.ICML 2026 · 14 citations
