On the choice of Perception Loss Function for Learned Video Compression
Sadaf Salehkalaibar, Buu Phan, Jun Chen, Wei Yu, Ashish Khisti
摘要
We study causal, low-latency, sequential video compression when the output is subjected to both a mean squared-error (MSE) distortion loss as well as a perception loss to target realism. Motivated by prior approaches, we consider two different perception loss functions (PLFs). The first, PLF-JD, considers the joint distribution (JD) of all the video frames up to the current one, while the second metric, PLF-FMD, considers the framewise marginal distributions (FMD) between the source and reconstruction. Using information theoretic analysis and deep-learning based experiments, we demonstrate that the choice of PLF can have a significant effect on the reconstruction, especially at low-bit rates. In particular, while the reconstruction based on PLF-JD can better preserve the temporal correlation across frames, it also imposes a significant penalty in distortion compared to PLF-FMD and further makes it more difficult to recover from errors made in the earlier output frames. Although the choice of PLF decisively affects reconstruction quality, we also demonstrate that it may not be essential to commit to a particular PLF during encoding and the choice of PLF can be delegated to the decoder. In particular, encoded representations generated by training a system to minimize the MSE (without requiring either PLF) can be near universal and can generate close to optimal reconstructions for either choice of PLF at the decoder. We validate our results using (one-shot) information-theoretic analysis, detailed study of the rate-distortion-perception tradeoff of the Gauss-Markov source model as well as deep-learning based experiments on moving MNIST, KTH and UVG datasets. * Equal Contribution Preprint. Under review.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper12
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 被引用 675 次
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte 等ICCV 2019 · 被引用 648 次
- Deep Contextual Video CompressionJiahao Li, Bin Li, Yan LuNeurIPS 2021 · 被引用 518 次
- ELF-VC: Efficient Learned Flexible-Rate Video CodingOren Rippel, Alexander G. Anderson, Kedar Tatwawadi, Sanjay Nair 等ICCV 2021 · 被引用 137 次
- Universal Rate-Distortion-Perception Representations for Lossy CompressionGeorge Zhang, Jingjing Qian, Jun Chen, Ashish KhistiNeurIPS 2021 · 被引用 108 次
相关 Paper
- On Perceptual Lossy Compression: The Cost of Perceptual Reconstruction and An Optimal Training FrameworkZeyu Yan, Fei Wen, Rendong Ying, Chao Ma 等ICML 2021 · 被引用 48 次
- Perceptual Kalman Filters: Online State Estimation under a Perfect Perceptual-Quality ConstraintDror Freirich, Tomer Michaeli, Ron MeirNeurIPS 2023 · 被引用 4 次
- Optimally Controllable Perceptual Lossy CompressionZeyu Yan, Fei Wen, Peilin LiuICML 2022 · 被引用 23 次
- Video Coding using Learned Latent GAN CompressionMustafa Shukor, Bharath Bhushan Damodaran, Xu Yao, Pierre HellierACM MM 2022 · 被引用 6 次
- Video Compression With Rate-Distortion AutoencodersAmirHossein Habibian, Ties van Rozendaal, Jakub M. Tomczak, Taco CohenICCV 2019 · 被引用 233 次
