Efficient Video Diffusion Models via Content-Frame Motion-Latent Decomposition
Sihyun Yu, Weili Nie, De-An Huang, Boyi Li, Jinwoo Shin, Anima Anandkumar
摘要
Video diffusion models have recently made great progress in generation quality, but are still limited by the high memory and computational requirements. This is because current video diffusion models often attempt to process high-dimensional videos directly. To tackle this issue, we propose content-motion latent diffusion model (CMD), a novel efficient extension of pretrained image diffusion models for video generation. Specifically, we propose an autoencoder that succinctly encodes a video as a combination of a content frame (like an image) and a low-dimensional motion latent representation. The former represents the common content, and the latter represents the underlying motion in the video, respectively. We generate the content frame by fine-tuning a pretrained image diffusion model, and we generate the motion latent representation by training a new lightweight diffusion model. A key innovation here is the design of a compact latent space that can directly utilizes a pretrained image diffusion model, which has not been done in previous latent video diffusion models. This leads to considerably better quality generation and reduced computational costs. For instance, CMD can sample a video 7.7 faster than prior approaches by generating a video of 5121024 resolution and length 16 in 3.1 seconds. Moreover, CMD achieves an FVD score of 212.7 on WebVid-10M, 27.3% better than the previous state-of-the-art of 292.4.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper17
- Amuse: Human-AI Collaborative Songwriting with Multimodal InspirationsYewon Kim, Sung-Ju Lee, Chris DonahueCHI 2025 · 被引用 35 次
- Token Perturbation Guidance for Diffusion ModelsJavad Rajabi, Soroush Mehraban, Seyedmorteza Sadat, Babak TaatiNeurIPS 2025 · 被引用 17 次
- Chain of World: World Model Thinking in Latent MotionFuxiang Yang, Donglin Di, Lulu Tang, Xuancheng Zhang 等CVPR 2026 · 被引用 11 次
- RealCam-I2V: Real-World Image-to-Video Generation with Interactive Complex Camera ControlTeng Li, Guangcong Zheng, Rui Jiang, Shuigen Zhan 等ICCV 2025 · 被引用 5 次
- GIViC: Generative Implicit Video CompressionGe Gao, Siyue Teng, Tianhao Peng, Fan Zhang 等ICCV 2025 · 被引用 4 次
它引用的顶会 Paper44
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Denoising Diffusion Implicit ModelsJiaming Song, Chenlin Meng, Stefano ErmonICLR 2021 · 被引用 11,743 次
相关 Paper
- Video Probabilistic Diffusion Models in Projected Latent SpaceSihyun Yu, Kihyuk Sohn, Subin Kim, Jinwoo ShinCVPR 2023
- VIDM: Video Implicit Diffusion ModelsKangfu Mei, Vishal M. PatelAAAI 2023 · 被引用 107 次
- REDUCIO! Generating 1K Video Within 16 Seconds Using Extremely Compressed Motion LatentsRui Tian, Qi Dai, Jianmin Bao, Kai Qiu 等ICCV 2025 · 被引用 2 次
- Conditional Image-to-Video Generation with Latent Flow Diffusion ModelsHaomiao Ni, Changhao Shi, Kai Li, Sharon X. Huang 等CVPR 2023
- WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelZongjian Li, Bin Lin, Yang Ye, Liuhan Chen 等CVPR 2025
