When does perceptual alignment benefit vision representations?
Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, Phillip Isola
Abstract
Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly weigh these attributes and thus make inferences misaligned with human perception. While vision representations have previously benefited from alignment in contexts like image generation, the utility of perceptually aligned representations in more general-purpose settings remains unclear. Here, we investigate how aligning vision model representations to human perceptual judgments impacts their usability across diverse computer vision tasks. We finetune state-of-the-art models on human similarity judgments for image triplets and evaluate them across standard vision benchmarks. We find that aligning models to perceptual judgments yields representations that improve upon the original backbones across many downstream tasks, including counting, segmentation, depth estimation, instance retrieval, and retrieval-augmented generation. In addition, we find that performance is widely preserved on other tasks, including specialized out-of-distribution domains such as in medical imaging and 3D environment frames. Our results suggest that injecting an inductive bias about human perceptual knowledge into vision models can contribute to better representations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2bf6ca1e-0786-4237-ae32-230ea3528c02Cited by top-tier papers6
- Tikzero: Zero-Shot Text-Guided Graphics Program SynthesisJonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka et al.ICCV 2025 · 24 citations
- PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & EditingAntonio Oroz, Matthias Nießner, Tobias KirschsteinCVPR 2026 · 6 citations
- MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided DiffusionDongyang Li, Kunpeng Xie, Mingyang Wu, Yiwei Kong et al.ICLR 2026
- ID-Sim: An Identity-Focused Similarity MetricJulia Chae, Nick Kolkin, Jui-Hsien Wang, Richard Zhang et al.CVPR 2026
- Human-Aligned Image Models Improve Visual Decoding from the BrainNona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Farzaneh Taleb et al.ICML 2025
Builds on31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni et al.NeurIPS 2020 · 19,162 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
Related papers
- DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic DataStephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai et al.NeurIPS 2023 · 413 citations
- Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language ModelsLuca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata, Matthias Bethge et al.ICML 2025
- Enriching ImageNet With Human Similarity Judgments and Psychological EmbeddingsBrett D. Roads, Bradley C. LoveCVPR 2021
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen et al.ICLR 2023 · 15 citations
- Learning What Helps: Task-Aligned Context Selection for Vision TasksJingyu Guo, Emir Konuk, Fredrik Strand, Christos Matsoukas et al.CVPR 2026
