VideoMAETok: Boosting Video Diffusion Models via Masked Autoencoders as Tokenizers
Zhan Tong, Tinne Tuytelaars
摘要
Latent diffusion models have become the dominant paradigm for video generation, making the video tokenizer a critical component. While most existing tokenizers are trained primarily for reconstruction, diffusion models are optimized to denoise heavily corrupted latents, which creates a mismatch between tokenizer training objectives and downstream generative learning. As a result, reconstruction metrics (e.g., rFVD) can be a poor proxy for generation quality (gFVD), and overly prioritizing reconstruction may even hinder diffusion training. We propose VideoMAE-Tok, a simple family of ViT-based video tokenizers trained explicitly as corruption-inversion models for latent video diffusion. VideoMAE-Tok builds on masked autoencoders: we (i) apply high-ratio token masking and encode only visible spatiotemporal tokens for efficiency, and (ii) corrupt latent tokens with interpolative Gaussian noise to better match the denoising nature of diffusion generators. Training under such corruption encourages latents that remain informative and well-conditioned for downstream denoising. Extensive experiments show that Video-MAETok consistently improves generation quality when paired with off-the-shelf diffusion models (SiT and LightningDiT), achieving state-ofthe-art gFVD on Kinetics-600 and UCF-101 while remaining compute-efficient. Code is available at https://github.com/yztongzhan/VideoMAETok.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- Diffusion Models Beat GANs on Image SynthesisPrafulla Dhariwal, Alexander Quinn NicholNeurIPS 2021 · 被引用 13,211 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scalable Diffusion Models with TransformersWilliam Peebles, Saining XieICCV 2023 · 被引用 5,568 次
相关 Paper
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
- Latent Denoising Makes Good TokenizersJiawei Yang, Tianhong Li, Lijie Fan, Yonglong Tian 等ICLR 2026 · 被引用 17 次
- Masked Autoencoders Are Effective Tokenizers for Diffusion ModelsHao Chen, Yujin Han, Fangyi Chen, Xiang Li 等ICML 2025
- OmniTokenizer: A Joint Image-Video Tokenizer for Visual GenerationJunke Wang, Yi Jiang, Zehuan Yuan, Bingyue Peng 等NeurIPS 2024 · 被引用 132 次
- Reconstruction vs. Generation: Taming Optimization Dilemma in Latent Diffusion ModelsJingfeng Yao, Bin Yang, Xinggang WangCVPR 2025
