VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAE
Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen
Abstract
Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying image VAEs to individual frames in isolation results in temporal inconsistencies and fails to compress temporal redundancy effectively. Existing works on Video VAEs compress temporal redundancy but struggle to handle videos with large motion effectively. They suffer from issues such as severe image blur and loss of detail in scenarios with large motion. In this paper, we present a powerful video VAE named VideoVAE+ that effectively reconstructs videos with large motion. First, we investigate two architecture choices and propose our simple yet effective architecture with better spatiotemporal joint modeling performance. Second, we propose to leverage the textual information in existing text-to-video datasets and incorporate text guidance during training. The textural guidance is optional during inference. We find that this design enhances the reconstruction quality and preservation of detail. Finally, our models achieve strong performance compared with various baseline approaches in both general videos and large motion videos, demonstrating its effectiveness on the challenging large motion scenarios.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6d74edac-6fc3-4e25-a9de-18200a6a43e7Cited by top-tier papers3
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with TransformersMinguk Kang, Suha KwakCVPR 2026 · 1 citation
- MoVie: Multimodal Video Compression with Text GuidanceJiaqi Hu, Haoji Hu, Heming Sun, Lianrui MuICML 2026
- VideoMAETok: Boosting Video Diffusion Models via Masked Autoencoders as TokenizersZhan Tong, Tinne TuytelaarsICML 2026
Builds on13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser et al.CVPR 2022 · 13,123 citations
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan et al.NeurIPS 2022 · 2,948 citations
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 1,550 citations
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang et al.ICLR 2024 · 1,493 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
Related papers
- High-Quality Joint Image and Video Tokenization with Causal VAEDawit Mureja Argaw, Xian Liu, Qinsheng Zhang, Joon Son Chung et al.ICLR 2025
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-ResolutionJunyang Chen, Jiangxin Dong, Long Sun, Yixin Yang et al.CVPR 2026 · 1 citation
- CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerZhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding et al.ICLR 2025
- UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingJianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo et al.ACM MM 2025 · 6 citations
- MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete DiffusionOnkar Kishor Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal et al.ICLR 2025
