Enriching ImageNet With Human Similarity Judgments and Psychological Embeddings
Brett D. Roads, Bradley C. Love
摘要
Advances in supervised learning approaches to object recognition flourished in part because of the availability of high-quality datasets and associated benchmarks. However, these benchmarks-such as ILSVRC-are relatively task-specific, focusing predominately on predicting class labels. We introduce a publicly-available dataset that embodies the task-general capabilities of human perception and reasoning. The Human Similarity Judgments extension to ImageNet (ImageNet-HSJ) is composed of a large set of human similarity judgments that supplements the existing ILSVRC validation set. The new dataset supports a range of task and performance metrics, including evaluation of unsupervised algorithms. We demonstrate two methods of assessment: using the similarity judgments directly and using a psychological embedding trained on the similarity judgments. This embedding space contains an order of magnitude more points (i.e., images) than previous efforts based on human judgments. We were able to scale to the full 50,000 image ILSVRC validation set through a selective sampling process that used variational Bayesian inference and model ensembles to sample aspects of the embedding space that were most uncertain. To demonstrate the utility of ImageNet-HSJ, we used the similarity ratings and the embedding space to evaluate how well several popular models conform to human similarity judgments. One finding is that the more complex models that perform better on task-specific benchmarks do not better conform to human semantic judgments. In addition to the human similarity judgments, pre-trained psychological embeddings and code for inferring variational embeddings are made publicly available. ImageNet-HSJ supports the appraisal of internal representations and the development of more humanlike models.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper10
- Harmonizing the object recognition strategies of deep neural networks with humansThomas Fel, Ivan F. Rodriguez Rodriguez, Drew Linsley, Thomas SerreNeurIPS 2022 · 被引用 111 次
- Probabilistic Contrastive Learning Recovers the Correct Aleatoric Uncertainty of Ambiguous InputsMichael Kirchhof, Enkelejda Kasneci, Seong Joon OhICML 2023 · 被引用 33 次
- VICE: Variational Interpretable Concept EmbeddingsLukas Muttenthaler, Charles Y. Zheng, Patrick McClure, Robert A. Vandermeulen 等NeurIPS 2022 · 被引用 29 次
- Human alignment of neural network representationsLukas Muttenthaler, Jonas Dippel, Lorenz Linhardt, Robert A. Vandermeulen 等ICLR 2023 · 被引用 15 次
- Evaluating alignment between humans and neural network representations in image-based learning tasksCan Demircan, Tankred Saanum, Leonardo Pettini, Marcel Binz 等NeurIPS 2024 · 被引用 11 次
它引用的顶会 Paper2
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Human Uncertainty Makes Classification More RobustJoshua C. Peterson, Ruairidh M. Battleday, Thomas L. Griffiths, Olga RussakovskyICCV 2019 · 被引用 362 次
相关 Paper
- VisRes Bench: On Evaluating the Visual Reasoning Capabilities of VLMsBrigitta Malagurski Törtei, Yasser Dahou, Ngoc Dung Huynh, Wamiq Reyaz Para 等CVPR 2026 · 被引用 3 次
- When does perceptual alignment benefit vision representations?Shobhita Sundaram, Stephanie Fu, Lukas Muttenthaler, Netanel Tamir 等NeurIPS 2024 · 被引用 24 次
- Mieb: Massive Image Embedding BenchmarkChenghao Xiao, Isaac Chung, Imene Kerboua, Jamie Stirling 等ICCV 2025
- GeneCIS: A Benchmark for General Conditional Image SimilaritySagar Vaze, Nicolas Carion, Ishan MisraCVPR 2023
- Beyond Counting: Evaluating Abstract and Emotional Reasoning in Vision-Language ModelsYuan Zhou, Yan Zhang, Jianlong Chang, Xin Gu 等AAAI 2026
