VDOT: Efficient Unified Video Creation via Optimal Transport Distillation
Yutong Wang, Haiyu Zhang, Tianfan Xue, Yu Qiao, Yaohui Wang, Chang Xu, Xinyuan Chen
Abstract
The rapid development of generative models has significantly advanced image and video applications. Among these, video creation, aimed at generating videos under various conditions, has gained substantial attention. However, existing video creation models either focus solely on a few specific conditions or suffer from excessively long generation times due to complex model inference, making them impractical for real-world applications. To mitigate these issues, we propose an efficient unified video creation model, named VDOT. Concretely, we model the training process with the distribution matching distillation (DMD) paradigm. Instead of using the Kullback-Leibler (KL) minimization, we additionally employ a novel computational optimal transport (OT) technique to optimize the discrepancy between the real and fake score distributions. The OT distance inherently imposes geometric constraints, mitigating potential zero-forcing or gradient collapse issues that may arise during KL-based distillation within the few-step generation scenario, and thus, enhances the efficiency and stability of the distillation process. Further, we integrate a discriminator to enable the model to perceive real video data, thereby enhancing the quality of generated videos. To support training unified video creation models, we propose a fully automated pipeline for video data annotation and filtering that accommodates multiple video creation tasks. Meanwhile, we curate a unified testing benchmark, UVCBench, to advance the field. Experiments demonstrate that our 4-step VDOT outperforms or matches the performance of other baselines with 100 denoising steps.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on38
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li et al.NeurIPS 2022 · 8,965 citations
- Adding Conditional Control to Text-to-Image Diffusion ModelsLvmin Zhang, Anyi Rao, Maneesh AgrawalaICCV 2023 · 6,759 citations
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari et al.ICML 2024 · 3,620 citations
- T2I-Adapter: Learning Adapters to Dig Out More Controllable Ability for Text-to-Image Diffusion ModelsChong Mou, Xintao Wang, Liangbin Xie, Yanze Wu et al.AAAI 2024 · 1,641 citations
Related papers
- Towards One-step Causal Video Generation via Adversarial Self-DistillationYongqi Yang, Huayang Huang, Xu Peng, Xiaobin Hu et al.ICLR 2026 · 17 citations
- OSV: One Step is Enough for High-Quality Image to Video GenerationXiaofeng Mao, Zhengkai Jiang, Fu-Yun Wang, Jiangning Zhang et al.CVPR 2025
- Adversarial Distribution Matching for Diffusion Distillation Towards Efficient Image and Video SynthesisYanzuo Lu, Yuxi Ren, Xin Xia, Shanchuan Lin et al.ICCV 2025 · 4 citations
- Transition Matching Distillation for Fast Video GenerationWeili Nie, Julius Berner, Nanye Ma, Chao Liu et al.CVPR 2026 · 24 citations
- AAD-1: Asymmetric Adversarial Distillation for One-Step Autoregressive Video GenerationHaobo Li, Yanhong Zeng, Yunhong Lu, Jiapeng Zhu et al.ICML 2026 · 1 citation
