Deep Hierarchical Video Compression
Ming Lu, Zhihao Duan, Fengqing Zhu, Zhan Ma
Abstract
Recently, probabilistic predictive coding that directly models the conditional distribution of latent features across successive frames for temporal redundancy removal has yielded promising results. Existing methods using a single-scale Variational AutoEncoder (VAE) must devise complex networks for conditional probability estimation in latent space, neglecting multiscale characteristics of video frames. Instead, this work proposes hierarchical probabilistic predictive coding, for which hierarchal VAEs are carefully designed to characterize multiscale latent features as a family of flexible priors and posteriors to predict the probabilities of future frames. Under such a hierarchical structure, lightweight networks are sufficient for prediction. The proposed method outperforms representative learned video compression models on common testing videos and demonstrates computational friendliness with much less memory footprint and faster encoding/decoding. Extensive experiments on adaptation to temporal patterns also indicate the better generalization of our hierarchical predictive mechanism. Furthermore, our solution is the first to enable progressive decoding that is favored in networked video applications with packet loss.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 353618dc-9273-420e-804d-a74f7a9a4e22Cited by top-tier papers9
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 20 citations
- Learned Image Transmission with Hierarchical Variational AutoencoderGuangyi Zhang, Hanlei Li, Yunlong Cai, Qiyu Hu et al.AAAI 2025 · 7 citations
- BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangACM MM 2025 · 3 citations
- Taming Hierarchical Image Coding Optimization: A Spectral Regularization PerspectiveWuyang Cong, Junqi Shi, Ming Lu, Xu Zhang et al.ICLR 2026
- Reinforced Rate Control for Neural Video Compression via Inter-Frame Rate-Distortion AwarenessWuyang Cong, Junqi Shi, Lizhong Wang, Weijing Shi et al.AAAI 2026
Builds on14
- A ConvNet for the 2020sZhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer et al.CVPR 2022 · 6,782 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma et al.CVPR 2022 · 363 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
Related papers
- Bit Prioritization in Variational Autoencoders via Progressive CodingRui Shu, Stefano ErmonICML 2022 · 9 citations
- Improved Conditional VRNNs for Video PredictionLluís Castrejón, Nicolas Ballas, Aaron C. CourvilleICCV 2019 · 177 citations
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- Hierarchical Quantized AutoencodersWill Williams, Sam Ringer, Tom Ash, David MacLeod et al.NeurIPS 2020 · 90 citations
- Variational Predictive Routing with Nested Subjective TimescalesAlexey Zakharov, Qinghai Guo, Zafeirios FountasICLR 2022 · 12 citations
