VIDM: Video Implicit Diffusion Models
Kangfu Mei, Vishal M. Patel
Abstract
Diffusion models have emerged as a powerful generative method for synthesizing high-quality and diverse set of images. In this paper, we propose a video generation method based on diffusion models, where the effects of motion are modeled in an implicit condition manner, i.e. one can sample plausible video motions according to the latent feature of frames. We improve the quality of the generated videos by proposing multiple strategies such as sampling space truncation, robustness penalty, and positional group normalization. Various experiments are conducted on datasets consisting of videos with different resolutions and different number of frames. Results show that the proposed method outperforms the state-of-the-art generative adversarial network-based methods by a significant margin in terms of FVD scores as well as perceptible visual quality.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 646599f7-26fa-462b-90f5-f23691430ac8Cited by top-tier papers47
- FIFO-Diffusion: Generating Infinite Videos from Text without TrainingJihwan Kim, Junoh Kang, Jinyoung Choi, Bohyung HanNeurIPS 2024 · 156 citations
- Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic TaskMaya Okawa, Ekdeep Singh Lubana, Robert P. Dick, Hidenori TanakaNeurIPS 2023 · 113 citations
- VillanDiffusion: A Unified Backdoor Attack Framework for Diffusion ModelsSheng-Yen Chou, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2023 · 101 citations
- Evaluation of Text-to-Video Generation Models: A Dynamics PerspectiveMingxiang Liao, Hannan Lu, Qixiang Ye, Wangmeng Zuo et al.NeurIPS 2024 · 89 citations
- Autodecoding Latent 3D Diffusion ModelsEvangelos Ntavelis, Aliaksandr Siarohin, Kyle Olszewski, Chaoyang Wang et al.NeurIPS 2023 · 65 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 13,211 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Improved Denoising Diffusion Probabilistic ModelsAlexander Quinn Nichol, Prafulla DhariwalICML 2021 · 5,234 citations
Related papers
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li et al.ICLR 2024 · 34 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Video Probabilistic Diffusion Models in Projected Latent SpaceSihyun Yu, Kihyuk Sohn, Subin Kim, Jinwoo ShinCVPR 2023
- Feature Prediction Diffusion Model for Video Anomaly DetectionCheng Yan, Shiyu Zhang, Yang Liu, Guansong Pang et al.ICCV 2023 · 76 citations
- STDD: Spatio-Temporal Dual Diffusion for Video GenerationShuaizhen Yao, Xiaoya Zhang, Xin Liu, Mengyi Liu et al.CVPR 2025
