Deep Learning in Latent Space for Video Prediction and Compression
Bowen Liu, Yu Chen, Shiyu Liu, Hun-Seok Kim
Abstract
Learning-based video compression has achieved substantial progress during recent years. The most influential approaches adopt deep neural networks (DNNs) to remove spatial and temporal redundancies by finding the appropriate lower-dimensional representations of frames in the video. We propose a novel DNN based framework that predicts and compresses video sequences in the latent vector space. The proposed method first learns the efficient lower-dimensional latent space representation of each video frame and then performs inter-frame prediction in that latent domain. The proposed latent domain compression of individual frames is obtained by a deep autoencoder trained with a generative adversarial network (GAN). To exploit the temporal correlation within the video frame sequence, we employ a convolutional long short-term memory (ConvLSTM) network to predict the latent vector representation of the future frame. We demonstrate our method with two applications; video compression and abnormal event detection that share the identical latent frame prediction network. The proposed method exhibits superior or competitive performance compared to the stateof-the-art algorithms specifically designed for either video compression or anomaly detection. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4fce599a-f3be-4dcb-ad40-e885a0732263Cited by top-tier papers9
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
- Comprehensive Regularization in a Bi-directional Predictive Network for Video Anomaly DetectionChengwei Chen, Yuan Xie, Shaohui Lin, Angela Yao et al.AAAI 2022 · 68 citations
- MMVP: Motion-Matrix-based Video PredictionYiqi Zhong, Luming Liang, Ilya Zharkov, Ulrich NeumannICCV 2023 · 39 citations
- Learning-Based Video Coding with Joint Deep Compression and EnhancementTiesong Zhao, Weize Feng, Hongji Zeng, Yiwen Xu et al.ACM MM 2022 · 24 citations
- Objects do not disappear: Video object detection by single-frame object location anticipationXin Liu, Fatemeh Karimi Nejadasl, Jan C. van Gemert, Olaf Booij et al.ICCV 2023 · 10 citations
Builds on7
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte et al.ICCV 2019 · 648 citations
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson et al.ICCV 2019 · 258 citations
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 207 citations
- Learning for Video Compression With Hierarchical Quality and Recurrent EnhancementRen Yang, Fabian Mentzer, Luc Van Gool, Radu TimofteCVPR 2020
Related papers
- M-LVC: Multiple Frames Prediction for Learned Video CompressionJianping Lin, Dong Liu, Houqiang Li, Feng WuCVPR 2020
- Hierarchical B-Frame Video Coding Using Two-Layer CANF Without Motion CodingDavid Alexandre, Hsueh-Ming Hang, Wen-Hsiao PengCVPR 2023
- High-Quality Joint Image and Video Tokenization with Causal VAEDawit Mureja Argaw, Xian Liu, Qinsheng Zhang, Joon Son Chung et al.ICLR 2025
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- FLAVC: Learned Video Compression with Feature Level AttentionChun Zhang, Heming Sun, Jiro KattoCVPR 2025
