Autoregressive Adversarial Post-Training for Real-Time Interactive Video Generation
Shanchuan Lin, Ceyuan Yang, Hao He, Jianwen Jiang, Yuxi Ren, Xin Xia, Yang Zhao, Xuefeng Xiao, Lu Jiang
Abstract
Existing large-scale video generation models are computationally intensive, preventing adoption in real-time and interactive applications. In this work, we propose autoregressive adversarial post-training (AAPT) to transform a pre-trained latent video diffusion model into a real-time, interactive video generator. Our model autoregressively generates a latent frame at a time using a single neural function evaluation (1NFE). The model can stream the result to the user in real time and receive interactive responses as controls to generate the next latent frame. Unlike existing approaches, our method explores adversarial training as an effective paradigm for autoregressive generation. This not only allows us to design an architecture that is more efficient for one-step generation while fully utilizing the KV cache, but also enables training the model in a student-forcing manner that proves to be effective in reducing error accumulation during long video generation. Our experiments demonstrate that our 8B model achieves real-time, 24fps, streaming video generation at 736×416 resolution on a single H100, or 1280×720 on 8×H100 up to a minute long (1440 frames).
In recent years, the field of visual content creation has been transformed by the rise of foundation models for video generation [4,78,69,44,95]. These models have enabled a wide range of powerful applications, including text-to-video generation, image-to-video synthesis, and controllable video creation conditioned on various multi-modal signals.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b4097a65-919e-4973-9522-124f7c531c7aCited by top-tier papers33
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan et al.ICLR 2026 · 215 citations
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang et al.ICLR 2026 · 181 citations
- WorldPlay: Towards Long-Term Geometric Consistency for Real-Time Interactive World ModelingWenqiang Sun, Haiyu Zhang, Haoyuan Wang, Junta Wu et al.ICML 2026 · 108 citations
- DreamDojo: A Real-Time Robot World Model from Large-Scale Human VideosShenyuan Gao, William Liang, Kaiyuan Zheng, Ayaan Malik et al.ICML 2026 · 96 citations
- Scaling Instruction-Based Video Editing with a High-Quality Synthetic DatasetQingyan Bai, Qiuyu Wang, Hao Ouyang, Yue Yu et al.CVPR 2026 · 79 citations
Builds on65
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 5,568 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
Related papers
- Diffusion Adversarial Post-Training for One-Step Video GenerationShanchuan Lin, Xin Xia, Yuxi Ren, Ceyuan Yang et al.ICML 2025
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao et al.ICLR 2026 · 241 citations
- SF-V: Single Forward Video Generation ModelZhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu et al.NeurIPS 2024 · 43 citations
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou et al.NeurIPS 2025 · 628 citations
- Flex-Forcing: Towards a Unified Autoregressive and Bidirectional Video Diffusion ModelXinyin Ma, Julius Berner, Chao Liu, Arash Vahdat et al.ICML 2026
