MMVC: Learned Multi-Mode Video Compression with Block-based Prediction Mode Selection and Density-Adaptive Entropy Coding
Bowen Liu, Yu Chen, Rakesh Chowdary Machineni, Shiyu Liu, Hun-Seok Kim
Abstract
Learning-based video compression has been extensively studied over the past years, but it still has limitations in adapting to various motion patterns and entropy models. In this paper, we propose multi-mode video compression (MMVC), a block wise mode ensemble deep video compression framework that selects the optimal mode for feature domain prediction adapting to different motion patterns. Proposed multi-modes include ConvLSTM-based feature domain prediction, optical flow conditioned feature domain prediction, and feature propagation to address a wide range of cases from static scenes without apparent motions to dynamic scenes with a moving camera. We partition the feature space into blocks for temporal prediction in spatial block-based representations. For entropy coding, we consider both dense and sparse post-quantization residual blocks, and apply optional run-length coding to sparse residuals to improve the compression rate. In this sense, our method uses a dual-mode entropy coding scheme guided by a binary density map, which offers significant rate reduction surpassing the extra cost of transmitting the binary selection map. We validate our scheme with some of the most popular benchmarking datasets. Compared with state-ofthe-art video compression schemes and standard codecs, our method yields better or competitive results measured with PSNR and MS-SSIM.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9630caa5-8f04-4c32-b642-2f903d3be2c4Cited by top-tier papers6
- Ultra-Fast Neural Video CompressionJiahao Li, Wenxuan Xie, Zhaoyang Jia, Bin Li et al.CVPR 2026 · 7 citations
- Neural Video Compression with In-Loop Contextual Filtering and Out-of-Loop Reconstruction EnhancementYaojun Wu, Chaoyi Lin, Yiming Wang, Semih Esenlik et al.ACM MM 2025 · 1 citation
- Neural Video Compression with Feature ModulationJiahao Li, Bin Li, Yan LuCVPR 2024
- Towards Practical Real-Time Neural Video CompressionZhaoyang Jia, Bin Li, Jiahao Li, Wenxuan Xie et al.CVPR 2025
- Implicit Motion FunctionYue Gao, Jiahao Li, Lei Chu, Yan LuCVPR 2024
Builds on12
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson et al.ICCV 2019 · 258 citations
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 233 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
Related papers
- Learned Video Compression via Joint Spatial-Temporal Correlation ExplorationHaojie Liu, Han Shen, Lichao Huang, Ming Lu et al.AAAI 2020 · 63 citations
- Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode PredictionZhihao Hu, Guo Lu, Jinyang Guo, Shan Liu et al.CVPR 2022 · 95 citations
- M-LVC: Multiple Frames Prediction for Learned Video CompressionJianping Lin, Dong Liu, Houqiang Li, Feng WuCVPR 2020
- FVC: A New Framework Towards Deep Video Compression in Feature SpaceZhihao Hu, Guo Lu, Dong XuCVPR 2021
- MoVie: Multimodal Video Compression with Text GuidanceJiaqi Hu, Haoji Hu, Heming Sun, Lianrui MuICML 2026
