WF-VAE: Enhancing Video VAE by Wavelet-Driven Energy Flow for Latent Video Diffusion Model
Zongjian Li, Bin Lin, Yang Ye, Liuhan Chen, Xinhua Cheng, Shenghai Yuan, Li Yuan
摘要
Video Variational Autoencoder (VAE) encodes videos into a low-dimensional latent space, becoming a key component of most Latent Video Diffusion Models (LVDMs) to reduce model training costs. However, as the resolution and duration of generated videos increase, the encoding cost of Video VAEs becomes a limiting bottleneck in training LVDMs. Moreover, the block-wise inference method adopted by most LVDMs can lead to discontinuities of latent space when processing long-duration videos. The key to addressing the computational bottleneck lies in decomposing videos into distinct components and efficiently encoding the critical information. Wavelet transform can decompose videos into multiple frequency-domain components and improve the efficiency significantly, we thus propose Wavelet Flow VAE (WF-VAE), an autoencoder that leverages multi-level wavelet transform to facilitate lowfrequency energy flow into latent representation. Furthermore, we introduce a method called Causal Cache, which maintains the integrity of latent space during block-wise inference. Compared to state-of-the-art video VAEs, WF-VAE demonstrates superior performance in both PSNR and LPIPS metrics, achieving 2× higher throughput and 4× lower memory consumption while maintaining competitive reconstruction quality. Our code and models are available at https://github.com/PKU-YuanGroup/WF-VAE.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- MAGREF: Masked Guidance for Any-Reference Video Generation with Subject DisentanglementYufan Deng, Yuanyang Yin, Xun Guo, Yizhi Wang 等ICLR 2026 · 被引用 20 次
- SIGMAN: Scaling 3D Human Gaussian Generation with Millions of AssetsYuhang Yang, Fengqi Liu, Yixing Lu, Qin Zhao 等ICCV 2025 · 被引用 5 次
- Progressive Growing of Video Tokenizers for Temporally Compact Latent SpacesAniruddha Mahapatra, Long Mai, David Bourgin, Yitian Zhang 等ICCV 2025 · 被引用 3 次
- MatPedia: A Universal Generative Foundation for High-Fidelity Material SynthesisDi Luo, Shuhui Yang, Mingxin Yang, Jiawei Lu 等CVPR 2026 · 被引用 3 次
- VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAEYazhou Xing, Yang Fei, Yingqing He, Jingye Chen 等ICCV 2025 · 被引用 2 次
它引用的顶会 Paper19
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Scaling Rectified Flow Transformers for High-Resolution Image SynthesisPatrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari 等ICML 2024 · 被引用 3,620 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
相关 Paper
- LiteVAE: Lightweight and Efficient Variational Autoencoders for Latent Diffusion ModelsSeyedmorteza Sadat, Jakob Buhmann, Derek Bradley, Otmar Hilliges 等NeurIPS 2024 · 被引用 29 次
- LeanVAE: An Ultra-Efficient Reconstruction VAE for Video Diffusion ModelsYu Cheng, Fajie YuanICCV 2025 · 被引用 1 次
- Video Probabilistic Diffusion Models in Projected Latent SpaceSihyun Yu, Kihyuk Sohn, Subin Kim, Jinwoo ShinCVPR 2023
- Efficient Video Diffusion Models via Content-Frame Motion-Latent DecompositionSihyun Yu, Weili Nie, De-An Huang, Boyi Li 等ICLR 2024 · 被引用 34 次
- Flash-VAED: Plug-and-Play VAE Decoders for Efficient Video GenerationLunjie Zhu, Yushi Huang, Xingtong Ge, Yufei Xue 等ICML 2026
