Streaming Autoregressive Video Generation via Diagonal Distillation
Jinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen, Ming-Hsuan Yang, Weiyang Liu
摘要
Large pretrained diffusion models have significantly enhanced the quality of generated videos, and yet their use in real-time streaming remains limited. Autoregressive models offer a natural framework for sequential frame synthesis but require heavy computation to achieve high fidelity.Diffusion distillation can compress these models into efficient few-step variants, but existing video distillation approaches largely adapt image-specific methods that neglect temporal dependencies. These techniques often excel in image generation but underperform in video synthesis, exhibiting reduced motion coherence, error accumulation over long sequences, and a latency-quality trade-off. We identify two factors that result in these limitations: insufficient utilization of temporal context during step reduction and implicit prediction of subsequent noise levels in next-chunk prediction (i.e., exposure bias). To address these issues, we propose Diagonal Distillation, which operates orthogonally to existing approaches and better exploits temporal information across both video chunks and denoising steps. Central to our approach is an asymmetric generation strategy: more steps early, fewer steps later.This design allows later chunks to inherit rich appearance information from thoroughly processed early chunks, while using partially denoised chunks as conditional inputs for subsequent synthesis. By aligning the implicit prediction of subsequent noise levels during chunk generation with the actual inference conditions, our approach mitigates error propagation and reduces oversaturation in long-range sequences. We further incorporate implicit optical flow modeling to preserve motion quality under strict step constraints.Our method generates a 5-second video in 2.61 seconds (up to 31 FPS), achieving a 277.3× speedup over the undistilled model. * Equal contribution with alphabetical order.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step GenerationYuyang You, Yongzhi Li, Jiahui Li, Yadong Mu 等CVPR 2026 · 被引用 7 次
- XYZFlow: Scaling Multidimensional Shortcut Flows for Efficient Generative ModelingJinxiu Liu, Xuanming Liu, Kangfu Mei, Yandong Wen 等ICML 2026
它引用的顶会 Paper29
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- ProlificDreamer: High-Fidelity and Diverse Text-to-3D Generation with Variational Score DistillationZhengyi Wang, Cheng Lu, Yikai Wang, Fan Bao 等NeurIPS 2023 · 被引用 1,498 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
相关 Paper
- Towards One-step Causal Video Generation via Adversarial Self-DistillationYongqi Yang, Huayang Huang, Xu Peng, Xiaobin Hu 等ICLR 2026 · 被引用 17 次
- InstantViR: Real-Time Video Inverse Problem Solver with Distilled Diffusion PriorWeimin Bai, Suzhe Xu, Yiwei Ren, Jinhua Hao 等CVPR 2026 · 被引用 3 次
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu 等CVPR 2026 · 被引用 24 次
- DisCa: Accelerating Video Diffusion Transformers with Distillation-Compatible Learnable Feature CachingChang Zou, Changlin Li, Songtao Liu, Zhao Zhong 等CVPR 2026 · 被引用 5 次
- Autoregressive Distillation of Diffusion TransformersYeongmin Kim, Sotiris Anagnostidis, Yuming Du, Edgar Schönfeld 等CVPR 2025
