MAGVIT: Masked Generative Video Transformer
Lijun Yu, Yong Cheng, Kihyuk Sohn, José Lezama, Han Zhang, Huiwen Chang, Alexander G. Hauptmann, Ming-Hsuan Yang, Yuan Hao, Irfan Essa, Lu Jiang
摘要
60x 250x (b) Efficiency (a) Quality 0 200 400 MAGVIT TATS (prior best) CCVS 76 -77% 332 386 0.0 45.0 90.0 MAGVIT TATS (prior best) Video Diffusion 89.3 +13% 79.3 57.0 30 60 90 MAGVIT RaMViD (prior best) NÜWA 62 -26% 84 87 0 10 20 MAGVIT Video Diffusion (prior best) RaMViD 9.9 -39% 16.2 16.5 UCF-101 CG FVD↓ BAIR FP FVD↓ Kinetics-600 FP FVD↓ UCF-101 CG IS↑ Estimated Relative Inference Runtime Inference Throughput At 128×128 native resolution MAGVIT-B 37 fps on 1x (c) Flexibility Class-conditional Generation (CG) Frame Prediction (FP) Frame Interpolation Outpainting Inpainting Squeezing Something 10 tasks in one model MAGVIT-L 65 fps on 1x And other tasks … GPU V100 TPU v4i Figure 1. Overview of the video generation quality, efficiency, and flexibility of the proposed MAGVIT model. (a) MAGVIT achieves the state-of-the-art FVD [61] and Inception Score (IS) [49] on two video generation tasks and three benchmarks, in comparison with prior best diffusion models (RaMViD [35], Video Diffusion [33]) and autoregressive models (CCVS [41], TATS [21], N ÜWA [70]). (b) It is two orders of magnitude faster than diffusion models and 60× faster than autoregressive models. (c) A single MAGVIT model accommodates different generation tasks, ranging from class-conditional generation to dynamic inpainting of a moving object.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper165
- Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale PredictionKeyu Tian, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 1,199 次
- Autoregressive Image Generation without Vector QuantizationTianhong Li, Yonglong Tian, He Li, Mingyang Deng 等NeurIPS 2024 · 被引用 758 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
- VideoPoet: A Large Language Model for Zero-Shot Video GenerationDan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama 等ICML 2024 · 被引用 464 次
- Pix2Video: Video Editing using Image DiffusionDuygu Ceylan, Chun-Hao Paul Huang, Niloy J. MitraICCV 2023 · 被引用 370 次
它引用的顶会 Paper23
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Zero-Shot Text-to-Image GenerationAditya Ramesh, Mikhail Pavlov, Gabriel Goh, Scott Gray 等ICML 2021 · 被引用 6,356 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
相关 Paper
- MCVD - Masked Conditional Video Diffusion for Prediction, Generation, and InterpolationVikram Voleti, Alexia Jolicoeur-Martineau, Chris PalNeurIPS 2022 · 被引用 434 次
- Autoregressive Video Generation without Vector QuantizationHaoge Deng, Ting Pan, Haiwen Diao, Zhengxiong Luo 等ICLR 2025
- BIVDiff: A Training-Free Framework for General-Purpose Video Synthesis via Bridging Image and Video Diffusion ModelsFengyuan Shi, Jiaxi Gu, Hang Xu, Songcen Xu 等CVPR 2024
- FlexiDiT: Your Diffusion Transformer Can Easily Generate High-Quality Samples with Less ComputeSotiris Anagnostidis, Gregor Bachmann, Yeongmin Kim, Jonas Kohler 等CVPR 2025
- FlowVid: Taming Imperfect Optical Flows for Consistent Video-to-Video SynthesisFeng Liang, Bichen Wu, Jialiang Wang, Licheng Yu 等CVPR 2024
