VLIC: Vision-Language Models As Perceptual Judges for Human-Aligned Image Compression
Kyle Sargent, Ruiqi Gao, Philipp Henzler, Charles Herrmann, Aleksander Holynski, Li Fei-Fei, Jiajun Wu, Jason Y. Zhang
Abstract
Evaluations of image compression performance which include human preferences have generally found that naive distortion functions such as MSE are insufficiently aligned to human perception.In order to align compression models to human perception, prior work has employed differentiable perceptual losses consisting of neural networks calibrated on large-scale datasets of human psycho-visual judgments. We show that, surprisingly, state-of-the-art vision-language models (VLMs) can replicate binary human two-alternative forced choice (2AFC) judgments zero-shot when asked to reason about the differences between pairs of images. Motivated to exploit the powerful zero-shot visual reasoning capabilities of VLMs, we propose Vision Language Models for Image Compression (VLIC), a diffusion-based image compression system designed to be post-trained with binary VLM judgments. VLIC leverages existing techniques for diffusion model post-training with preferences, rather than distilling the VLM judgments into a separate perceptual loss network. We show that calibrating this system on VLM judgments produces competitive or state-of-the-art performance on human-aligned visual compression depending on the dataset, according to perceptual metrics and large-scale user studies. We additionally conduct an extensive analysis of the VLM-based reward design and training procedure and share important insights.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d95fb2a3-6e3d-4485-9824-9d55497c4d22Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning et al.NeurIPS 2023 · 10,924 citations
- Training Diffusion Models with Reinforcement LearningKevin Black, Michael Janner, Yilun Du, Ilya Kostrikov et al.ICLR 2024 · 816 citations
- High-Fidelity Generative Image CompressionFabian Mentzer, George Toderici, Michael Tschannen, Eirikur AgustssonNeurIPS 2020 · 675 citations
- Language Model Beats Diffusion - Tokenizer is key to visual generationLijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari et al.ICLR 2024 · 609 citations
Related papers
- Beyond VLM-Based Rewards: Diffusion-Native Latent Reward ModelingGongye Liu, Bo Yang, Zhi Yida, Zhizhou Zhong et al.ICML 2026 · 3 citations
- Diffusion Model as a Noise-Aware Latent Reward Model for Step-Level Preference OptimizationTao Zhang, Cheng Da, Kun Ding, Huan Yang et al.NeurIPS 2025 · 38 citations
- Training-Free Diffusion Model Alignment with Sampling DemonsPo-Hung Yeh, Kuang-Huei Lee, Jun-Cheng ChenICLR 2025
- Diff-ICMH: Harmonizing Machine and Human Vision in Image Compression with Generative PriorRuoyu Feng, Yunpeng Qi, Jinming Liu, Yixin Gao et al.NeurIPS 2025 · 5 citations
- SPRO: Improving Image Generation via Self-PlayRitika Jha, Aanisha Bhattacharyya, Yaman Singla, Rajiv Ratn Shah et al.NeurIPS 2025
