Towards One-step Causal Video Generation via Adversarial Self-Distillation
Yongqi Yang, Huayang Huang, Xu Peng, Xiaobin Hu, Donghao Luo, Jiangning Zhang, Chengjie Wang, Yu Wu
摘要
Recent hybrid video generation models combine autoregressive temporal dynamics with diffusion-based spatial denoising, but their sequential, iterative nature leads to error accumulation and long inference times. In this work, we propose a distillation-based framework for efficient causal video generation that enables high-quality synthesis with extremely limited denoising steps. Our approach builds upon the Distribution Matching Distillation (DMD) framework and proposes a novel Adversarial Self-Distillation (ASD) strategy, which aligns the outputs of the student model's n-step denoising process with its (n+1)-step version at the distribution level. This design provides smoother supervision by bridging small intra-student gaps and more informative guidance by combining teacher knowledge with locally consistent student behavior, substantially improving training stability and generation quality in extremely few-step scenarios (e.g., 1-2 steps). In addition, we present a First-Frame Enhancement (FFE) strategy, which allocates more denoising steps to the initial frames to mitigate error propagation while applying larger skipping steps to later frames. Extensive experiments on VBench demonstrate that our method surpasses state-of-the-art approaches in both one-step and two-step video generation. Notably, our framework produces a single distilled model that flexibly supports multiple inference-step settings, eliminating the need for repeated re-distillation and enabling efficient, high-quality video synthesis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Causality in Video Diffusers is Separable from DenoisingXingjian Bai, Guande He, Zhengqi Li, Eli Shechtman 等CVPR 2026 · 被引用 3 次
- Restoring Initial Noise Sensitivity in Text-to-Image Distillation through Geometric Alignmenthuayang Huang, Ruoyu Wang, Jinhui Zhao, Wei Deng 等ICML 2026
- Causal Forcing: Autoregressive Diffusion Distillation Done Right for High-Quality Real-Time Interactive Video GenerationHongzhou Zhu, Min Zhao, Guande He, Hang Su 等ICML 2026
它引用的顶会 Paper38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Adaptive Video Distillation: Mitigating Oversaturation and Temporal Collapse in Few-Step GenerationYuyang You, Yongzhi Li, Jiahui Li, Yadong Mu 等CVPR 2026 · 被引用 7 次
- AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video GenerationHaobo Li, Yanhong Zeng, Yunhong Lu, Jiapeng Zhu 等ICML 2026 · 被引用 1 次
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu 等CVPR 2026 · 被引用 24 次
- DUO-VSR: Dual-Stream Distillation for One-Step Video Super-ResolutionZhengyao Lv, Menghan Xia, Xintao Wang, Kwan-Yee K. WongCVPR 2026 · 被引用 4 次
- DOLLAR: Few-Step Video Generation Via Distillation and Latent Reward OptimizationZihan Ding, Chi Jin, Difan Liu, Haitian Zheng 等ICCV 2025 · 被引用 3 次
