VideoVAE+: Large Motion Video Autoencoding with Cross-Modal Video VAE
Yazhou Xing, Yang Fei, Yingqing He, Jingye Chen, Jiaxin Xie, Xiaowei Chi, Qifeng Chen
摘要
Learning a robust video Variational Autoencoder (VAE) is essential for reducing video redundancy and facilitating efficient video generation. Directly applying image VAEs to individual frames in isolation results in temporal inconsistencies and fails to compress temporal redundancy effectively. Existing works on Video VAEs compress temporal redundancy but struggle to handle videos with large motion effectively. They suffer from issues such as severe image blur and loss of detail in scenarios with large motion. In this paper, we present a powerful video VAE named VideoVAE+ that effectively reconstructs videos with large motion. First, we investigate two architecture choices and propose our simple yet effective architecture with better spatiotemporal joint modeling performance. Second, we propose to leverage the textual information in existing text-to-video datasets and incorporate text guidance during training. The textural guidance is optional during inference. We find that this design enhances the reconstruction quality and preservation of detail. Finally, our models achieve strong performance compared with various baseline approaches in both general videos and large motion videos, demonstrating its effectiveness on the challenging large motion scenarios.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- FlashDecoder: Real-Time Latent-to-Pixel Streaming Decoder with TransformersMinguk Kang, Suha KwakCVPR 2026 · 被引用 1 次
- MoVie: Multimodal Video Compression with Text GuidanceJiaqi Hu, Haoji Hu, Heming Sun, Lianrui MuICML 2026
- VideoMAETok: Boosting Video Diffusion Models via Masked Autoencoders as TokenizersZhan Tong, Tinne TuytelaarsICML 2026
它引用的顶会 Paper13
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Video Diffusion ModelsJonathan Ho, Tim Salimans, Alexey A. Gritsenko, William Chan 等NeurIPS 2022 · 被引用 2,948 次
- Frozen in Time: A Joint Video and Image Encoder for End-to-End RetrievalMax Bain, Arsha Nagrani, Gül Varol, Andrew ZissermanICCV 2021 · 被引用 1,550 次
- AnimateDiff: Animate Your Personalized Text-to-Image Diffusion Models without Specific TuningYuwei Guo, Ceyuan Yang, Anyi Rao, Zhengyang Liang 等ICLR 2024 · 被引用 1,493 次
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari 等ICLR 2024 · 被引用 609 次
相关 Paper
- High-Quality Joint Image and Video Tokenization with Causal VAEDawit Mureja Argaw, Xian Liu, Qinsheng Zhang, Joon Son Chung 等ICLR 2025
- STCDiT: Spatio-Temporally Consistent Diffusion Transformer for High-Quality Video Super-ResolutionJunyang Chen, Jiangxin Dong, Long Sun, Yixin Yang 等CVPR 2026 · 被引用 1 次
- CogVideoX: Text-to-Video Diffusion Models with An Expert TransformerZhuoyi Yang, Jiayan Teng, Wendi Zheng, Ming Ding 等ICLR 2025
- UniEdit: A Unified Tuning-Free Framework for Video Motion and Appearance EditingJianhong Bai, Tianyu He, Yuchi Wang, Junliang Guo 等ACM MM 2025 · 被引用 6 次
- MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete DiffusionOnkar Kishor Susladkar, Jishu Sen Gupta, Chirag Sehgal, Sparsh Mittal 等ICLR 2025
