Improving neural network representations using human similarity judgments
Lukas Muttenthaler, Lorenz Linhardt, Jonas Dippel, Robert A. Vandermeulen, Katherine L. Hermann, Andrew K. Lampinen, Simon Kornblith
Abstract
Deep neural networks have reached human-level performance on many computer vision tasks. However, the objectives used to train these networks enforce only that similar images are embedded at similar locations in the representation space, and do not directly constrain the global structure of the resulting space. Here, we explore the impact of supervising this global structure by linearly aligning it with human similarity judgments. We find that a naive approach leads to large changes in local representational structure that harm downstream performance. Thus, we propose a novel method that aligns the global structure of representations while preserving their local structure. This global-local transform considerably improves accuracy across a variety of few-shot learning and anomaly detection tasks. Our results indicate that human visual representations are globally organized in a way that facilitates learning from few examples, and incorporating this global structure into neural network representations improves performance on downstream tasks. * Work done as part of the Google Research Collabs Programme.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d11f31ce-9d7c-4417-953e-4f4b28889124Cited by top-tier papers10
- A Spectral Theory of Neural Prediction and AlignmentAbdulkadir Canatar, Jenelle Feather, Albert J. Wakhloo, SueYeon ChungNeurIPS 2023 · 29 citations
- When does perceptual alignment benefit vision representations?Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir et al.NeurIPS 2024 · 24 citations
- Evaluating alignment between humans and neural network representations in image-based learning tasksCan Demircan, Tankred Saanum, Leonardo Pettini, Marcel Binz et al.NeurIPS 2024 · 11 citations
- Learning Human-like Representations to Enable Learning Human ValuesAndrea Wynn, Ilia Sucholutsky, Tom GriffithsNeurIPS 2024 · 11 citations
- Set Learning for Accurate and Calibrated ModelsLukas Muttenthaler, Robert A. Vandermeulen, Qiuyi Zhang, Thomas Unterthiner et al.ICLR 2024 · 4 citations
Builds on22
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Emerging Properties in Self-Supervised Vision TransformersMathilde Caron, Hugo Touvron, Ishan Misra, Hervé Jégou et al.ICCV 2021 · 8,921 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Alignment with human representations supports robust few-shot learningIlia Sucholutsky, Tom GriffithsNeurIPS 2023 · 41 citations
- On Equivariant and Invariant Learning of Object Landmark RepresentationsZezhou Cheng, Jong-Chyi Su, Subhransu MajiICCV 2021 · 18 citations
- Exploring Complementary Strengths of Invariant and Equivariant Representations for Few-Shot LearningMamshad Nayeem Rizve, Salman H. Khan, Fahad Shahbaz Khan, Mubarak ShahCVPR 2021
- Channel Importance Matters in Few-Shot Image ClassificationXu Luo, Jing Xu, Zenglin XuICML 2022 · 57 citations
- A Hierarchical Transformation-Discriminating Generative Model for Few Shot Anomaly DetectionShelly Sheynin, Sagie Benaim, Lior WolfICCV 2021 · 106 citations
