Mobile Video Diffusion
Haitam Ben Yahia, Denis Korzhenkov, Ioannis Lelekas, Amir Ghodrati, Amirhossein Habibian
摘要
Video diffusion models have achieved impressive realism and controllability but are limited by high computational demands, restricting their use on mobile devices. This paper introduces the first mobile-optimized video diffusion model. Starting from a spatio-temporal UNet from Stable Video Diffusion (SVD), we reduce memory and computational cost by reducing the frame resolution, incorporating multi-scale temporal representations, and introducing two novel pruning schema to reduce the number of channels and temporal blocks. Furthermore, we employ adversarial finetuning to reduce the denoising to a single step. Our model, coined as MobileVD, is 523× more efficient (1817.2 vs. 4.34 TFLOPs) with a slight quality drop (FVD 149 vs. 171), generating latents for a 14 × 512 × 256 px clip in 1.7 seconds on a Xiaomi-14 Pro. Our results are available at https://qualcomm-ai-research.github.io/ mobile-video-diffusion/
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Attention Surgery: An Efficient Recipe to Linearize Your Video Diffusion TransformerMohsen Ghafoorian, Denis Korzhenkov, Amirhossein HabibianCVPR 2026 · 被引用 13 次
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile DevicesYa Zou, Jingfeng Yao, Siyuan Yu, Shuai Zhang 等AAAI 2026 · 被引用 8 次
- ReHyAt: Recurrent Hybrid Attention for Video Diffusion TransformersMohsen Ghafoorian, Amirhossein HabibianCVPR 2026 · 被引用 5 次
- PyramidalWan: On Making Pretrained Video Model Pyramidal for Efficient InferenceDenis Korzhenkov, Adil Karjauv, Animesh Karnewar, Mohsen Ghafoorian 等CVPR 2026 · 被引用 3 次
- V.I.P.: Iterative Online Preference Distillation for Efficient Video Diffusion ModelsJisoo Kim, Wooseok Seo, Junwan Kim, Seungho Park 等ICCV 2025 · 被引用 1 次
它引用的顶会 Paper29
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- DPM-Solver: A Fast ODE Solver for Diffusion Probabilistic Model Sampling in Around 10 StepsCheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen 等NeurIPS 2022 · 被引用 2,653 次
- Consistency ModelsYang Song, Prafulla Dhariwal, Mark Chen, Ilya SutskeverICML 2023 · 被引用 1,720 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Text2Video-Zero: Text-to-Image Diffusion Models are Zero-Shot Video GeneratorsLevon Khachatryan, Andranik Movsisyan, Vahram Tadevosyan, Roberto Henschel 等ICCV 2023 · 被引用 800 次
相关 Paper
- SF-V: Single Forward Video Generation ModelZhixing Zhang, Yanyu Li, Yushu Wu, Yanwu Xu 等NeurIPS 2024 · 被引用 43 次
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li 等ICLR 2024 · 被引用 34 次
- NanoSD: Edge Efficient Foundation Model for Real Time Image RestorationSubhajit Sanyal, Srinivas Soumitri Miriyala, Akshay Janardan Bankar, Manjunath Arveti 等CVPR 2026 · 被引用 3 次
- Training-Free Adaptive Diffusion with Bounded Difference Approximation StrategyHancheng Ye, Jiakang Yuan, Renqiu Xia, Xiangchao Yan 等NeurIPS 2024 · 被引用 21 次
- Efficiency-optimized Video Diffusion ModelsZijun Deng, Xiangteng He, Yuxin PengACM MM 2023 · 被引用 6 次
