Perceptual Video Compression with Neural Wrapping
Muhammad Umar Karim Khan, Aaron Chadha, Mohammad Ashraful Anam, Yiannis Andreopoulos
Abstract
Standard video codecs are rate-distortion optimization machines, where distortion is typically quantified using PSNR versus the source. However, it is now widely accepted that increasing PSNR does not necessarily translate to better visual quality. In this paper, a better balance between perception and fidelity is targeted, in order to provide for significant rate savings over state-of-the-art standards-based video codecs. Specifically, pre-and postprocessing neural networks are proposed that enhance the coding efficiency of standard video codecs when benchmarked with an array of well-established perceptual quality scores. These "neural wrapper" elements are end-to-end trained with a neural codec module serving as a differentiable proxy for standard video codecs. The codec proxy is jointly optimized with the pre-and post components via a novel two-phase pretraining strategy and end-to-end iterative refinement with stop-gradient. This allows the neural pre-and postprocessor to learn to embed, remove and recover information in a codec-aware manner, thus improving its rate-quality performance. A single neural-wrapper model is thereby established and used for the entire ratequality curve without needing any downscaling or upscaling. The trained model is tested with the AV1 and VVC standard codecs via an array of well-established objective quality scores (SSIM, MS-SSIM, VMAF, AVQT), as well as mean opinion scores (MOS) derived from ITU-T P.910 subjective testing. Experimental results show that the proposed approach improves all quality scores, with -18.5% average Bjontegaard Delta-rate (BD-rate) saving over all objective scores and MOS improvement over both standard codecs. This illustrates the significant potential of neural wrapper components over standards-based video coding. Post-processor (O) Pre-processor (P) Distortion Loss Rate Loss Muti-Stage Alignment Losses Codec Model (M) Codec (C) Losses to optimize wrappers Losses to optimize Codec Model
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on19
- BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and AlignmentKelvin C. K. Chan, Shangchen Zhou, Xiangyu Xu, Chen Change LoyCVPR 2022 · 522 citations
- Progressive Fusion Video Super-Resolution Network via Exploiting Non-Local Spatio-Temporal CorrelationsPeng Yi, Zhongyuan Wang, Kui Jiang, Junjun Jiang et al.ICCV 2019 · 309 citations
- Hybrid Spatial-Temporal Entropy Modelling for Neural Video CompressionJiahao Li, Bin Li, Yan LuACM MM 2022 · 202 citations
- Spatio-Temporal Deformable Convolution for Compressed Video Quality EnhancementJianing Deng, Li Wang, Shiliang Pu, Cheng ZhuoAAAI 2020 · 168 citations
- VCT: A Video Compression TransformerFabian Mentzer, George Toderici, David Minnen, Sergi Caelles et al.NeurIPS 2022 · 155 citations
Related papers
- NVRC: Neural Video Representation CompressionHo Man Kwan, Ge Gao, Fan Zhang, Andrew Gower et al.NeurIPS 2024 · 44 citations
- Perceptual Neural Video Compression with Color Separation and Rank Chainxiongzhuang liang, Chuanbo Tang, Zhuoyuan Li, Li Li et al.CVPR 2026
- Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image CompressionChuqin Zhou, Guo Lu, Jiangchuan Li, Xiangyu Chen et al.AAAI 2025 · 3 citations
- Video Compression with Entropy-Constrained Neural RepresentationsCarlos Gomes, Roberto Azevedo, Christopher SchroersCVPR 2023
- Efficient Adaptation of Neural Network Filter for Video CompressionYat Hong Lam, Alireza Zare, Francesco Cricri, Jani Lainema et al.ACM MM 2020 · 35 citations
