From Slow Bidirectional to Fast Autoregressive Video Diffusion Models
Tianwei Yin, Qiang Zhang, Richard Zhang, William T. Freeman, Frédo Durand, Eli Shechtman, Xun Huang
摘要
https://causvid.github.io/ "Macro shot of a man wearing an antique diving helmet with dark glass and a jetpack walking on the veins of a leaf. Realistic style" Bidirectional teacher Causal student Latency (gen. full 128-frame video) 219s Asymmetric distillation with DMD Initial Latency 1.3s On-the-fly generation 9.4 FPS … Figure 1. Traditional bidirectional diffusion models (top) deliver high-quality outputs but suffer from significant latency, taking 219 seconds to generate a 128-frame video. Users must wait for the entire sequence to complete before viewing any results. In contrast, we distill the bidirectional diffusion model into a few-step autoregressive generator (bottom), dramatically reducing computational overhead. Our model (CausVid) achieves an initial latency of only 1.3 seconds, after which frames are generated continuously in a streaming fashion at approximately 9.4 FPS, facilitating interactive workflows for video content creation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper114
- Self Forcing: Bridging the Train-Test Gap in Autoregressive Video DiffusionXun Huang, Zhengqi Li, Guande He, Mingyuan Zhou 等NeurIPS 2025 · 被引用 628 次
- LongLive: Real-time Interactive Long Video GenerationShuai Yang, Wei Huang, Ruihang Chu, Yicheng Xiao 等ICLR 2026 · 被引用 241 次
- Rolling Forcing: Autoregressive Long Video Diffusion in Real TimeKunhao Liu, Wenbo Hu, Jiale Xu, Ying Shan 等ICLR 2026 · 被引用 215 次
- Self-Forcing++: Towards Minute-Scale High-Quality Video GenerationJiaxing Cui, Jie Wu, Ming Li, Tao Yang 等ICLR 2026 · 被引用 181 次
- Video World Models with Long-term Spatial MemoryTong Wu, Shuai Yang, Ryan Po, Yinghao Xu 等NeurIPS 2025 · 被引用 145 次
它引用的顶会 Paper61
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationHongzhou Zhu, Min Zhao, Guande He, Hang Su 等ICML 2026
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao 等CVPR 2026 · 被引用 3 次
- MotionStream: Real-Time Video Generation with Interactive Motion ControlsJoonghyuk Shin, Zhengqi Li, Richard Zhang, Jun-Yan Zhu 等ICLR 2026 · 被引用 79 次
- Streaming Autoregressive Video Generation via Diagonal DistillationJinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen 等ICLR 2026 · 被引用 16 次
- Ca2-VDM: Efficient Autoregressive Video Diffusion Model with Causal Generation and Cache SharingKaifeng Gao, Jiaxin Shi, Hanwang Zhang, Chunping Wang 等ICML 2025
