Improved Video VAE for Latent Video Diffusion Model
Pingyu Wu, Kai Zhu, Yu Liu, Liming Zhao, Wei Zhai, Yang Cao, Zheng-Jun Zha
Abstract
Variational Autoencoder (VAE) aims to compress pixel data into low-dimensional latent space, playing an important role in OpenAI's Sora and other latent video diffusion generation models. While most existing video VAEs inflate a pre-trained image VAE into the 3D causal structure for temporal-spatial compression, this paper presents two astonishing findings: (1) The initialization from a well-trained image VAE with the same latent dimensions is not an optimal scheme. (2) The adoption of causal reasoning leads to unequal information interactions and unbalanced performance between frames. To alleviate these problems, we propose a keyframe-based temporal compression (KTC) architecture and a group causal convolution (GCConv) module to further improve video VAE (IV-VAE). Specifically, the KTC architecture divides the latent space into two branches, in which one half completely inherits the compression prior of keyframes from a lower-dimension image VAE while the other half involves temporal compression through 3D group causal convolution, reducing temporal-spatial conflicts and accelerating the convergence speed of video VAE. The GC-Conv in the above 3D half uses standard convolution within each frame group to ensure inter-frame equivalence, and employs causal logical padding between groups to maintain flexibility in processing variable frame video. Extensive experiments on five benchmarks demonstrate the SOTA video reconstruction and generation abilities of our IV-VAE.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers17
- Routing Matters in MoE: Scaling Diffusion Transformers with Explicit Routing GuidanceYujie Wei, Shiwei Zhang, Hangjie Yuan, Yujin Han et al.ICLR 2026 · 26 citations
- PMQ-VE: Progressive Multi-Frame Quantization for Video EnhancementZhanfeng Feng, Long Peng, Xin Di, Yong Guo et al.NeurIPS 2025 · 17 citations
- Towards Sequence Modeling Alignment between Tokenizer and Autoregressive ModelPingyu Wu, Kai Zhu, Yu Liu, Longxiang Tang et al.ICLR 2026 · 16 citations
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li et al.CVPR 2026 · 7 citations
- YOSE: You Only Select Essential Tokens for Efficient DiT-based Video Object RemovalChenyang Wu, Lina Lei, Fan Li, Chunle Guo et al.CVPR 2026 · 4 citations
Builds on12
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- SDXL: Improving Latent Diffusion Models for High-Resolution Image SynthesisDustin Podell, Zion English, Kyle Lacey, Andreas Blattmann et al.ICLR 2024 · 4,569 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
- CV-VAE: A Compatible Video VAE for Latent Generative Video ModelsSijie Zhao, Yong Zhang, Xiaodong Cun, Shaoshu Yang et al.NeurIPS 2024 · 82 citations
- Upscale-A-Video: Temporal-Consistent Diffusion Model for Real-World Video Super-ResolutionShangchen Zhou, Peiqing Yang, Jianyi Wang, Yihang Luo et al.CVPR 2024 · 52 citations
Related papers
- High-Quality Joint Image and Video Tokenization with Causal VAEDawit Mureja Argaw, Xian Liu, Qinsheng Zhang, Joon Son Chung et al.ICLR 2025
- WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion ModelZongjian Li, Bin Lin, Yang Ye, Liuhan Chen et al.CVPR 2025
- Real-Time Generation of Streamable Talking Portrait Video with Reference-Guided Deep Compression VAEsSicheng Xu, Yu Deng, Shoukang Hu, Yichuan Wang et al.CVPR 2026 · 1 citation
- Generative Latent Diffusion for Efficient Spatiotemporal Data ReductionXiao Li, Liangji Zhu, Anand Rangarajan, Sanjay RankaSC 2025 · 1 citation
- VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAEYazhou Xing, Yang Fei, Yingqing He, Jingye Chen et al.ICCV 2025 · 2 citations
