MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video Coding
Huu-Tai Phung, Zong-Lin Gao, Yi-Chen Yao, Kuan-Wei Ho, Yi-Hsin Chen, Yu-Hsiang Lin, Alessandro Gnutti, Wen-Hsiao Peng
Abstract
This work, termed MH-LVC, presents a multi-hypothesis temporal prediction scheme that employs long- and shortterm reference frames in a conditional residual video coding framework. Recent temporal context mining approaches to conditional video coding offer superior coding performance. However, the need to store and access a large amount of implicit contextual information extracted from past decoded frames in decoding a video frame poses a challenge due to excessive memory access. Our MH-LVC overcomes this issue by storing multiple long- and shortterm reference frames but limiting the number of reference frames used at a time for temporal prediction to two. Our decoded frame buffer management allows the encoder to flexibly utilize the long-term key frames to mitigate temporal cascading errors and the short-term reference frames to minimize prediction errors. Moreover, our buffering scheme enables the temporal prediction structure to be adapted to individual input videos. While this flexibility is common in traditional video codecs, it has not been fully explored for learned video codecs. Extensive experiments show that the proposed method outperforms VTM-17.0 under the low-delay B configuration in terms of PSNR-RGB across commonly used test datasets, and performs comparably to the state-of-the-art learned codecs (e.g. DCVCFM) while requiring less decoded frame buffer and similar decoding time. The source code of MH-LVC is available at https://github.com/NYCU-MAPL/MHLVC.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on9
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 518 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- ELF-VC: Efficient Learned Flexible-Rate Video CodingOren Rippel, Alexander G. Anderson, Kedar Tatwawadi, Sanjay Nair et al.ICCV 2021 · 137 citations
- Neural Video Compression with Diverse ContextsJiahao Li, Bin Li, Yan LuCVPR 2023
- M-LVC: Multiple Frames Prediction for Learned Video CompressionJianping Lin, Dong Liu, Houqiang Li, Feng WuCVPR 2020
Related papers
- ECVC: Exploiting Non-Local Correlations in Multiple Frames for Contextual Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangCVPR 2025
- BiECVC: Gated Diversification of Bidirectional Contexts for Learned Video CompressionWei Jiang, Junru Li, Kai Zhang, Li ZhangACM MM 2025 · 3 citations
- Neural Video Compression with Context ModulationChuanbo Tang, Zhuoyuan Li, Yifan Bian, Li Li et al.CVPR 2025
- HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video CodingYi-Hsin Chen, Yi-Chen Yao, Kuan-Wei Ho, Chun-Hung Wu et al.ICCV 2025 · 3 citations
- Neural Video Compression with Reference HierarchyChuanbo Tang, Zhuoyuan Li, Li Li, Dong Liu et al.AAAI 2026
