RIFLEx: A Free Lunch for Length Extrapolation in Video Diffusion Transformers
Min Zhao, Guande He, Yixiao Chen, Hongzhou Zhu, Chongxuan Li, Jun Zhu
摘要
Recent advancements in video generation have enabled models to synthesize high-quality, minutelong videos. However, generating even longer videos with temporal coherence remains a major challenge and existing length extrapolation methods lead to temporal repetition or motion deceleration. In this work, we systematically analyze the role of frequency components in positional embeddings and identify an intrinsic frequency that primarily governs extrapolation behavior. Based on this insight, we propose RIFLEx, a minimal yet effective approach that reduces the intrinsic frequency to suppress repetition while preserving motion consistency, without requiring any additional modifications. RIFLEx offers a true free lunch-achieving high-quality 2× extrapolation on state-of-the-art video diffusion transformers in a completely training-free manner. Moreover, it enhances quality and enables 3× extrapolation by minimal fine-tuning without long videos. Project page and codes: https://riflex-video.github.io/ .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper28
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao 等ICLR 2026 · 被引用 241 次
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang 等ICLR 2026 · 被引用 181 次
- Sparse VideoGen2: Accelerate Video Generation with Sparse Attention via Semantic-Aware PermutationShuo Yang, Haocheng Xi, Yilong Zhao, Muyang Li 等NeurIPS 2025 · 被引用 114 次
- Reward Forcing: Efficient Streaming Video Generation with Rewarded Distribution Matching DistillationYunhong Lu, Yanhong Zeng, Haobo Li, Hao Ouyang 等CVPR 2026 · 被引用 77 次
- Radial Attention: 𝒪(n log n) Sparse Attention with Energy Decay for Long Video GenerationXingyang Li, Muyang Li, Tianle Cai, Haocheng Xi 等NeurIPS 2025 · 被引用 66 次
它引用的顶会 Paper24
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Score-Based Generative Modeling through Stochastic Differential EquationsYang Song, Jascha Sohl-Dickstein, Diederik P. Kingma, Abhishek Kumar 等ICLR 2021 · 被引用 1,270 次
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 被引用 1,168 次
- YaRN: Efficient Context Window Extension of Large Language ModelsBowen Peng, Jeffrey Quesnelle, Honglu Fan, Enrico ShippoleICLR 2024 · 被引用 508 次
相关 Paper
- UltraViCo: Breaking Extrapolation Limits in Video Diffusion TransformersMin Zhao, Hongzhou Zhu, Yingze Wang, Bokai Yan 等ICLR 2026 · 被引用 14 次
- FreeLong: Training-Free Long Video Generation with SpectralBlend Temporal AttentionYu Lu, Yuanzhi Liang, Linchao Zhu, Yi YangNeurIPS 2024 · 被引用 101 次
- LongDiff: Training-Free Long Video Generation in One GoZhuoling Li, Hossein Rahmani, Qiuhong Ke, Jun LiuCVPR 2025
- Dlfr-Gen: Diffusion-Based Video Generation With Dynamic Latent Frame RateZhihang Yuan, Rui Xie, Yuzhang Shang, Hanling Zhang 等ICCV 2025 · 被引用 1 次
- Free-Lunch Long Video Generation via Layer-Adaptive O.O.D CorrectionJiahao Tian, Chenxi Song, Wei Cheng, Chi ZhangCVPR 2026 · 被引用 3 次
