MAGVIT: Masked Generative Video Transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, Lu Jiang
Abstract
60x 250x (b) Efficiency (a) Quality 0 200 400 MAGVIT TATS (prior best) CCVS 76 -77% 332 386 0.0 45.0 90.0 MAGVIT TATS (prior best) Video Diffusion 89.3 +13% 79.3 57.0 30 60 90 MAGVIT RaMViD (prior best) NÜWA 62 -26% 84 87 0 10 20 MAGVIT Video Diffusion (prior best) RaMViD 9.9 -39% 16.2 16.5 UCF-101 CG FVD↓ BAIR FP FVD↓ Kinetics-600 FP FVD↓ UCF-101 CG IS↑ Estimated Relative Inference Runtime Inference Throughput At 128×128 native resolution MAGVIT-B 37 fps on 1x (c) Flexibility Class-conditional Generation (CG) Frame Prediction (FP) Frame Interpolation Outpainting Inpainting Squeezing Something 10 tasks in one model MAGVIT-L 65 fps on 1x And other tasks … GPU V100 TPU v4i Figure 1. Overview of the video generation quality, efficiency, and flexibility of the proposed MAGVIT model. (a) MAGVIT achieves the state-of-the-art FVD [61] and Inception Score (IS) [49] on two video generation tasks and three benchmarks, in comparison with prior best diffusion models (RaMViD [35], Video Diffusion [33]) and autoregressive models (CCVS [41], TATS [21], N ÜWA [70]). (b) It is two orders of magnitude faster than diffusion models and 60× faster than autoregressive models. (c) A single MAGVIT model accommodates different generation tasks, ranging from class-conditional generation to dynamic inpainting of a moving object.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers165
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng et al.NeurIPS 2024 · 1,199 citations
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng et al.NeurIPS 2024 · 758 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama et al.ICML 2024 · 464 citations
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 370 citations
Builds on23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray et al.ICML 2021 · 6,356 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
Related papers
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 434 citations
- Autoregressive Video Generation without Vector QuantizationHaoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo et al.ICLR 2025
- BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion ModelsFengyuan Shi, Jiaxi Gu, Hang Xu, Songcen Xu et al.CVPR 2024
- FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less ComputeSotiris Anagnostidis, Gregor Bachmann, Yeongmin Kim, Jonas Kohler et al.CVPR 2025
- FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisFeng Liang, Bichen Wu, Jialiang Wang, Licheng Yu et al.CVPR 2024
