ICML2026

Deep Forcing: Training-Free Long Video Generation with Deep Sink and Participative Compression

Jung Yi, Wooseok Jang, Paul Cho, Jisu Nam, Heeji Yoon, Seungryong Kim

摘要

Recent advances in autoregressive video diffusion have enabled real-time frame streaming, however, existing methods still suffer from visual error accumulation including visual fidelity and motion degradation over long-horizon. To address these challenges, we introduce Deep Forcing, a training-free extension of autoregressive video diffusion models that stabilizes long video generation through two complementary mechanisms. Deep Sink preserves approximately half of the sliding context window as persistent sink tokens and realigns their temporal RoPE phases to the current timeline, thereby maintaining global context during extended rollouts. Participative Compression performs importance-aware KV cache pruning, retaining only tokens that actively participate in recent attention while removing redundant or degraded history, effectively mitigating error accumulation under out-of-distribution lengths. Together, these components enable over 12× length extrapolation (e.g., 5s-trained → 60s+) without sacrificing inference speed, while improving visual fidelity and motion dynamics compared to prior methods. Our results demonstrate that Deep Forcing can achieve performance comparable to state-of-the-art training-based methods trained specifically for long video generation.