HyTIP: Hybrid Temporal Information Propagation for Masked Conditional Residual Video Coding
Yi-Hsin Chen, Yi-Chen Yao, Kuan-Wei Ho, Chun-Hung Wu, Huu-Tai Phung, Martin Benjak, Jörn Ostermann, Wen-Hsiao Peng
摘要
Most frame-based learned video codecs can be interpreted as recurrent neural networks (RNNs) propagating reference information along the temporal dimension. This work revisits the limitations of the current approaches from an RNN perspective. The output-recurrence methods, which propagate decoded frames, are intuitive but impose dual constraints on the output decoded frames, leading to suboptimal rate-distortion performance. In contrast, the hidden-to-hidden connection approaches, which propagate latent features within the RNN, offer greater flexibility but require large buffer sizes. To address these issues, we propose HyTIP, a learned video coding framework that combines both mechanisms. Our hybrid buffering strategy uses explicit decoded frames and a small number of implicit latent features to achieve competitive coding performance. Experimental results show that our HyTIP outperforms the sole use of either output-recurrence or hidden-to-hidden approaches. Furthermore, it achieves comparable performance to state-of-the-art methods but with a much smaller buffer size, and outperforms VTM 17.0 (Low-delay B) in terms of PSNR-RGB and MS-SSIM-RGB. The source code of HyTIP is available at https://github.com/NYCU-MAPL/HyTIP.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper13
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson 等ICCV 2019 · 被引用 258 次
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 被引用 202 次
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles 等NeurIPS 2022 · 被引用 155 次
- Coarse-To-Fine Deep Video Coding with Hyperprior-Guided Mode PredictionZhihao Hu, Guo Lu, Jinyang Guo, Shan Liu 等CVPR 2022 · 被引用 95 次
相关 Paper
- MH-LVC: Multi-Hypothesis Temporal Prediction for Learned Conditional Residual Video CodingHuu-Tai Phung, Zong-Lin Gao, Yi-Chen Yao, Kuan-Wei Ho 等ICCV 2025 · 被引用 2 次
- M-LVC: Multiple Frames Prediction for Learned Video CompressionJianping Lin, Dong Liu, Houqiang Li, Feng WuCVPR 2020
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 被引用 20 次
- HNeRV: A Hybrid Neural Representation for VideosHao Chen, Matthew Gwilliam, Ser-Nam Lim, Abhinav ShrivastavaCVPR 2023
- Neural Inter-Frame Compression for Video CodingAbdelaziz Djelouah, Joaquim Campos, Simone Schaub-Meyer, Christopher SchroersICCV 2019 · 被引用 207 次
