Pragmatic Image Compression for Human-in-the-Loop Decision-Making
Siddharth Reddy, Anca D. Dragan, Sergey Levine
Abstract
Standard lossy image compression algorithms aim to preserve an image's appearance, while minimizing the number of bits needed to transmit it. However, the amount of information actually needed by a user for downstream tasks -- e.g., deciding which product to click on in a shopping website -- is likely much lower. To achieve this lower bitrate, we would ideally only transmit the visual features that drive user behavior, while discarding details irrelevant to the user's decisions. We approach this problem by training a compression model through human-in-the-loop learning as the user performs tasks with the compressed images. The key insight is to train the model to produce a compressed image that induces the user to take the same action that they would have taken had they seen the original image. To approximate the loss function for this model, we train a discriminator that tries to distinguish whether a user's action was taken in response to the compressed image or the original. We evaluate our method through experiments with human participants on four tasks: reading handwritten digits, verifying photos of faces, browsing an online shopping catalogue, and playing a car racing video game. The results show that our method learns to match the user's actions with and without compression at lower bitrates than baseline methods, and adapts the compression model to the user's behavior: it preserves the digit number and randomizes handwriting style in the digit reading task, preserves hats and eyeglasses while randomizing faces in the photo verification task, preserves the perceived price of an item while randomizing its color and background in the online shopping task, and preserves upcoming bends in the road in the car racing game.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 73961554-7797-41d4-9772-1cbc2de2cbc5Cited by top-tier papers2
- Remember the Past: Distilling Datasets into Addressable Memories for Neural NetworksZhiwei Deng, Olga RussakovskyNeurIPS 2022 · 140 citations
- Perceptual Group Tokenizer: Building Perception with Iterative GroupingZhiwei Deng, Ting Chen, Yang LiICLR 2024 · 4 citations
Builds on6
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- NVAE: A Deep Hierarchical Variational AutoencoderArash Vahdat, Jan KautzNeurIPS 2020 · 1,141 citations
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 675 citations
- Learning Representations by Humans, for HumansSophie Hilgard, Nir Rosenfeld, Mahzarin R. Banaji, Jack Cao et al.ICML 2021 · 33 citations
- Encoding in Style: A StyleGAN Encoder for Image-to-Image TranslationElad Richardson, Yuval Alaluf, Or Patashnik, Yotam Nitzan et al.CVPR 2021
Related papers
- Generative Adversarial Networks for Extreme Learned Image CompressionEirikur Agustsson, Michael Tschannen, Fabian Mentzer, Radu Timofte et al.ICCV 2019 · 648 citations
- Discernible Image CompressionZhaohui Yang, Yunhe Wang, Chang Xu, Peng Du et al.ACM MM 2020 · 25 citations
- Lossy Compression for Lossless PredictionYann Dubois, Benjamin Bloem-Reddy, Karen Ullrich, Chris J. MaddisonNeurIPS 2021 · 82 citations
- Multi-Modality Deep Network for Extreme Learned Image CompressionXuhao Jiang, Weimin Tan, Tian Tan, Bo Yan et al.AAAI 2023 · 27 citations
- VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image CompressionKyle Sargent, Ruiqi Gao, Philipp Henzler, Charles Herrmann et al.CVPR 2026 · 1 citation
