LIVE: Long-horizon Interactive Video World Modeling
Junchao Huang, Ziyang Ye, Xinting Hu, Tianyu He, Guiyu Zhang, Shaoshuai Shi, Jiang Bian, Li Jiang
摘要
Autoregressive video world models predict future visual observations conditioned on actions. While effective over short horizons, these models often struggle with long-horizon generation, as small prediction errors accumulate over time. Prior methods alleviate this by introducing pre-trained teacher models and sequence-level distribution matching, which incur additional computational cost and fail to prevent error propagation beyond the training horizon. In this work, we propose LIVE, a Long-horizon Interactive Video world modEl that enforces bounded error accumulation via a novel cycle-consistency objective, thereby eliminating the need for teacher-based distillation. Specifically, LIVE first performs a forward rollout from ground-truth frames and then applies a reverse generation process to reconstruct the initial state. The diffusion loss is subsequently computed on the reconstructed terminal state, providing an explicit constraint on long-horizon error propagation. Moreover, we provide a unified view that encompasses different approaches and introduce progressive training curriculum to stabilize training. Experiments demonstrate that LIVE achieves state-of-the-art performance on long-horizon benchmarks, generating stable, high-quality videos far beyond training rollout lengths.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper23
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Diffusion Forcing: Next-token Prediction Meets Full-Sequence DiffusionBoyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz 等NeurIPS 2024 · 被引用 751 次
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 等ICML 2024 · 被引用 464 次
相关 Paper
- BAgger: Backwards Aggregation for Mitigating Drift in Autoregressive Video Diffusion ModelsRyan Po, Eric Ryan Chan, Changan Chen, Gordon WetzsteinCVPR 2026 · 被引用 18 次
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang 等ICLR 2026 · 被引用 181 次
- Learning World Models for Interactive Video GenerationTaiye Chen, Xun Hu, Zihan Ding, Chi JinNeurIPS 2025 · 被引用 37 次
- Endless World: Real-Time 3D-Aware Long Video GenerationKe Zhang, Jiacong Xu, Yiqun Mei, Vishal M. PatelCVPR 2026 · 被引用 4 次
- Accelerating Autoregressive Video Diffusion via History-Guided Cache and Residual CorrectionKepan Nan, Wangbo Zhao, Penghao Zhou, Jun Li 等CVPR 2026
