Enriching ImageNet With Human Similarity Judgments and Psychological Embeddings
Brett D. Roads, Bradley C. Love
Abstract
Advances in supervised learning approaches to object recognition flourished in part because of the availability of high-quality datasets and associated benchmarks. However, these benchmarks-such as ILSVRC-are relatively task-specific, focusing predominately on predicting class labels. We introduce a publicly-available dataset that embodies the task-general capabilities of human perception and reasoning. The Human Similarity Judgments extension to ImageNet (ImageNet-HSJ) is composed of a large set of human similarity judgments that supplements the existing ILSVRC validation set. The new dataset supports a range of task and performance metrics, including evaluation of unsupervised algorithms. We demonstrate two methods of assessment: using the similarity judgments directly and using a psychological embedding trained on the similarity judgments. This embedding space contains an order of magnitude more points (i.e., images) than previous efforts based on human judgments. We were able to scale to the full 50,000 image ILSVRC validation set through a selective sampling process that used variational Bayesian inference and model ensembles to sample aspects of the embedding space that were most uncertain. To demonstrate the utility of ImageNet-HSJ, we used the similarity ratings and the embedding space to evaluate how well several popular models conform to human similarity judgments. One finding is that the more complex models that perform better on task-specific benchmarks do not better conform to human semantic judgments. In addition to the human similarity judgments, pre-trained psychological embeddings and code for inferring variational embeddings are made publicly available. ImageNet-HSJ supports the appraisal of internal representations and the development of more humanlike models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 03936df1-986f-48f1-b97e-1f189f5dfd84Cited by top-tier papers10
- Harmonizing the object recognition strategies of deep neural networks with humansThomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas SerreNeurIPS 2022 · 111 citations
- Probabilistic Contrastive Learning Recovers the Correct Aleatoric Uncertainty of Ambiguous InputsMichael Kirchhof, Enkelejda Kasneci, Seong Joon OhICML 2023 · 33 citations
- VICE: Variational Interpretable Concept EmbeddingsLukas Muttenthaler, Charles Y. Zheng, Patrick McClure, Robert A. Vandermeulen et al.NeurIPS 2022 · 29 citations
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen et al.ICLR 2023 · 15 citations
- Evaluating alignment between humans and neural network representations in image-based learning tasksCan Demircan, Tankred Saanum, Leonardo Pettini, Marcel Binz et al.NeurIPS 2024 · 11 citations
Builds on2
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 362 citations
Related papers
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para et al.CVPR 2026 · 3 citations
- When does perceptual alignment benefit vision representations?Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir et al.NeurIPS 2024 · 24 citations
- Mieb: Massive Image Embedding BenchmarkChenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling et al.ICCV 2025
- GeneCIS: A Benchmark for General Conditional Image SimilaritySagar Vaze, Nicolas Carion, Ishan MisraCVPR 2023
- Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language ModelsYuan Zhou, Yan Zhang, Jianlong Chang, Xin Gu et al.AAAI 2026
