What Matters in Practical Learned Image Compression
Kedar Tatwawadi, Parisa Rahimzadeh, Zhanghao Sun, Zhiqi Chen, Ziyun Yang, Sanjay Nair, Divija Hasteer, Oren Rippel
摘要
One of the major differentiators unlocked by learned codecs relative to their hard-coded traditional counterparts is their ability to be optimized directly to appeal to the human visual system. Despite this potential, a perceptual yet practical image codec is yet to be proposed.
In this work, we aim to close this gap. We conduct a comprehensive study of the key modeling choices that govern the design of a practical learned image codec, jointly optimized for perceptual quality and runtime -including within the ablations several novel techniques. We then perform performance-aware neural architecture search over millions of backbone configurations to identify models that achieve the target on-device runtime while maximizing compression performance as captured by perceptual metrics.
We combine the various optimizations to construct a new codec that achieves a significantly improved tradeoff between speed and perceptual quality. Based on rigorous subjective user studies, it provides 2.3-3× bitrate savings against AV1, AV2, VVC, ECM and JPEG-AI, and 20-40% bitrate savings against the best learned codec alternatives. At the same time, on an iPhone 17 Pro Max, it encodes 12MP images as fast as 230ms, and decodes them in 150ms -faster than most top ML-based codecs run on a V100 GPU.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 被引用 675 次
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte 等ICCV 2019 · 被引用 648 次
- ELIC: Efficient Learned Image Compression with Unevenly Grouped Space-Channel Contextual Adaptive CodingDailan He, Ziming Yang, Weikun Peng, Rui Ma 等CVPR 2022 · 被引用 363 次
- Lossy Image Compression with Conditional Diffusion ModelsRuihan Yang, Stephan MandtNeurIPS 2023 · 被引用 268 次
- ELF-VC: Efficient Learned Flexible-Rate Video CodingOren Rippel, Alexander G. Anderson, Kedar Tatwawadi, Sanjay Nair 等ICCV 2021 · 被引用 137 次
相关 Paper
- Controllable Distortion-Perception Tradeoff Through Latent Diffusion for Neural Image CompressionChuqin Zhou, Guo Lu, Jiangchuan Li, Xiangyu Chen 等AAAI 2025 · 被引用 3 次
- Efficient Learned Image Compression without Entropy CodingHao Cao, Wenqi Guo, Zhijin Qin, Jungong HanICML 2026
- Learned Video CompressionOren Rippel, Sanjay Nair, Carissa Lew, Steve Branson 等ICCV 2019 · 被引用 258 次
- Slimmable Compressive Autoencoders for Practical Neural Image CompressionFei Yang, Luis Herranz, Yongmei Cheng, Mikhail G. MozerovCVPR 2021
- PNVC: Towards Practical INR-based Video CompressionGe Gao, Ho Man Kwan, Fan Zhang, David BullAAAI 2025 · 被引用 20 次
