Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression
Jung Yi, Wooseok Jang, Paul Cho, Jisu Nam, Heeji Yoon, Seungryong Kim
摘要
Recent advances in autoregressive video diffusion have enabled real-time frame streaming, however, existing methods still suffer from visual error accumulation including visual fidelity and motion degradation over long-horizon. To address these challenges, we introduce Deep Forcing, a training-free extension of autoregressive video diffusion models that stabilizes long video generation through two complementary mechanisms. Deep Sink preserves approximately half of the sliding context window as persistent sink tokens and realigns their temporal RoPE phases to the current timeline, thereby maintaining global context during extended rollouts. Participative Compression performs importance-aware KV cache pruning, retaining only tokens that actively participate in recent attention while removing redundant or degraded history, effectively mitigating error accumulation under out-of-distribution lengths. Together, these components enable over 12× length extrapolation (e.g., 5s-trained → 60s+) without sacrificing inference speed, while improving visual fidelity and motion dynamics compared to prior methods. Our results demonstrate that Deep Forcing can achieve performance comparable to state-of-the-art training-based methods trained specifically for long video generation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Mode Seeking meets Mean Seeking for Fast Long Video GenerationShengqu Cai, Weili Nie, Chao Liu, Julius Berner 等ICML 2026 · 被引用 9 次
- EMFormer: Efficient Multi-Scale Transformer for Accumulative Context Weather Forecastinghao chen, Tao Han, Jie ZHANG, Song Guo 等ICML 2026
- Yume1.5: A Text-Controlled Interactive World Generation ModelXiaofeng Mao, Zhen Li, Chuanhao Li, Xiaojie Xu 等CVPR 2026
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationHongzhou Zhu, Min Zhao, Guande He, Hang Su 等ICML 2026
它引用的顶会 Paper18
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Efficient Streaming Language Models with Attention SinksGuangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han 等ICLR 2024 · 被引用 1,714 次
- H2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language ModelsZhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen 等NeurIPS 2023 · 被引用 1,003 次
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz 等NeurIPS 2024 · 被引用 751 次
- InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and GenerationYi Wang, Yinan He, Yizhuo Li, Kunchang Li 等ICLR 2024 · 被引用 467 次
相关 Paper
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan 等ICLR 2026 · 被引用 215 次
- LoL: Longer than Longer, Scaling Video Generation to HourJustin Cui, Jie Wu, Ming Li, Tao Yang 等CVPR 2026 · 被引用 30 次
- Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion ModelXinyin Ma, Julius Berner, Chao Liu, Arash Vahdat 等ICML 2026
- Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual CorrectionKepan Nan, Wangbo Zhao, Penghao Zhou, Jun Li 等CVPR 2026
