When does perceptual alignment benefit vision representations?
Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir, Lucy Chai, Simon Kornblith, Trevor Darrell, Phillip Isola
摘要
Humans judge perceptual similarity according to diverse visual attributes, including scene layout, subject location, and camera pose. Existing vision models understand a wide range of semantic abstractions but improperly weigh these attributes and thus make inferences misaligned with human perception. While vision representations have previously benefited from alignment in contexts like image generation, the utility of perceptually aligned representations in more general-purpose settings remains unclear. Here, we investigate how aligning vision model representations to human perceptual judgments impacts their usability across diverse computer vision tasks. We finetune state-of-the-art models on human similarity judgments for image triplets and evaluate them across standard vision benchmarks. We find that aligning models to perceptual judgments yields representations that improve upon the original backbones across many downstream tasks, including counting, segmentation, depth estimation, instance retrieval, and retrieval-augmented generation. In addition, we find that performance is widely preserved on other tasks, including specialized out-of-distribution domains such as in medical imaging and 3D environment frames. Our results suggest that injecting an inductive bias about human perceptual knowledge into vision models can contribute to better representations.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- Tikzero: Zero-Shot Text-Guided Graphics Program SynthesisJonas Belouadi, Eddy Ilg, Margret Keuper, Hideki Tanaka 等ICCV 2025 · 被引用 24 次
- PercHead: Perceptual Head Model for Single-Image 3D Head Reconstruction & EditingAntonio Oroz, Matthias Nießner, Tobias KirschsteinCVPR 2026 · 被引用 6 次
- MindPilot: Closed-loop Visual Stimulation Optimization for Brain Modulation with EEG-guided DiffusionDongyang Li, Kunpeng Xie, Mingyang Wu, Yiwei Kong 等ICLR 2026
- ID-Sim: An Identity-Focused Similarity MetricJulia Chae, Nick Kolkin, Jui-Hsien Wang, Richard Zhang 等CVPR 2026
- Human-Aligned Image Models Improve Visual Decoding from the BrainNona Rajabi, Antônio H. Ribeiro, Miguel Vasco, Farzaneh Taleb 等ICML 2025
它引用的顶会 Paper31
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Retrieval-Augmented Generation for Knowledge-Intensive NLP TasksPatrick Lewis, Ethan Perez, Aleksandra Piktus, Fabio Petroni 等NeurIPS 2020 · 被引用 19,162 次
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou 等ICCV 2021 · 被引用 8,921 次
相关 Paper
- DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic DataStephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai 等NeurIPS 2023 · 被引用 413 次
- Testing the Limits of Fine-Tuning for Improving Visual Cognition in Vision Language ModelsLuca M. Schulze Buschoff, Konstantinos Voudouris, Elif Akata, Matthias Bethge 等ICML 2025
- Enriching ImageNet With Human Similarity Judgments and Psychological EmbeddingsBrett D. Roads, Bradley C. LoveCVPR 2021
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen 等ICLR 2023 · 被引用 15 次
- Learning What Helps: Task-Aligned Context Selection for Vision TasksJingyu Guo, Emir Konuk, Fredrik Strand, Christos Matsoukas 等CVPR 2026
