CV-VAE: A Compatible Video VAE for Latent Generative Video Models
Sijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang, Muyao Niu, Xiaoyu Li, Wenbo Hu, Ying Shan
摘要
Spatio-temporal compression of videos, utilizing networks such as Variational Autoencoders (VAE), plays a crucial role in OpenAI's SORA and numerous other video generative models. For instance, many LLM-like video models learn the distribution of discrete tokens derived from 3D VAEs within the VQVAE framework, while most diffusion-based video models capture the distribution of continuous latent extracted by 2D VAEs without quantization. The temporal compression is simply realized by uniform frame sampling which results in unsmooth motion between consecutive frames. Currently, there lacks of a commonly used continuous video (3D) VAE for latent diffusion-based video models in the research community. Moreover, since current diffusion-based approaches are often implemented using pre-trained text-to-image (T2I) models, directly training a video VAE without considering the compatibility with existing T2I models will result in a latent space gap between them, which will take huge computational resources for training to bridge the gap even with the T2I models as initialization. To address this issue, we propose a method for training a video VAE of latent video models, namely CV-VAE, whose latent space is compatible with that of a given image VAE, e.g., image VAE of Stable Diffusion (SD). The compatibility is achieved by the proposed novel latent space regularization, which involves formulating a regularization loss using the image VAE. Benefiting from the latent space compatibility, video models can be trained seamlessly from pre-trained T2I or video models in a truly spatio-temporally compressed latent space, rather than simply sampling video frames at equal intervals. With our CV-VAE, existing video models can generate four times more frames with minimal finetuning. Extensive experiments are conducted to demonstrate the effectiveness of the proposed video VAE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper31
- Geometry-aware 4D Video Generation for Robot ManipulationZeyi Liu, Shuang Li, Eric Cousineau, Siyuan Feng 等ICLR 2026 · 被引用 28 次
- Turbo-VAED: Fast and Stable Transfer of Video-VAEs to Mobile DevicesYa Zou, Jingfeng Yao, Siyuan Yu, Shuai Zhang 等AAAI 2026 · 被引用 8 次
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li 等CVPR 2026 · 被引用 7 次
- SceneTok: A Compressed, Diffusable Token Space for 3D ScenesMohammad Asim, Christopher Wewer, Jan LenssenCVPR 2026 · 被引用 6 次
- MatPedia: A Universal Generative Foundation for High-Fidelity Material SynthesisDi Luo, Shuhui Yang, Mingxin Yang, Jiawei Lu 等CVPR 2026 · 被引用 3 次
相关 Paper
- Improved Video VAE for Latent Video Diffusion ModelPingyu Wu, Kai Zhu, Yu Liu, Liming Zhao 等CVPR 2025
- ColorDiffuser: Video Colorization with Pretrained Text-to-Image Diffusion ModelsHanyuan Liu, Minshan Xie, Jinbo Xing, Chengze Li 等ACM MM 2025 · 被引用 2 次
- REGEN: Learning Compact Video Embedding with (Re-)Generative DecoderYitian Zhang, Long Mai, Aniruddha Mahapatra, David Bourgin 等ICCV 2025
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li 等ICLR 2024 · 被引用 34 次
- WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelZongjian Li, Bin Lin, Yang Ye, Liuhan Chen 等CVPR 2025
