Words are all you need? Language as an approximation for human similarity judgments
Raja Marjieh, Pol van Rijn, Ilia Sucholutsky, Theodore R. Sumers, Harin Lee, Thomas L. Griffiths, Nori Jacoby
Abstract
Human similarity judgments are a powerful supervision signal for machine learning applications based on techniques such as contrastive learning, information retrieval, and model alignment, but classical methods for collecting human similarity judgments are too expensive to be used at scale. Recent methods propose using pre-trained deep neural networks (DNNs) to approximate human similarity, but pre-trained DNNs may not be available for certain domains (e.g., medical images, low-resource languages) and their performance in approximating human similarity has not been extensively tested. We conducted an evaluation of 611 pre-trained models across three domains -- images, audio, video -- and found that there is a large gap in performance between human similarity judgments and pre-trained DNNs. To address this gap, we propose a new class of similarity approximation methods based on language. To collect the language data required by these new methods, we also developed and validated a novel adaptive tag collection pipeline. We find that our proposed language-based methods are significantly cheaper, in the number of human judgments, than classical methods, but still improve performance over the DNN-based methods. Finally, we also develop `stacked' methods that combine language embeddings with DNN embeddings, and find that these consistently provide the best approximations for human similarity across all three of our modalities. Based on the results of this comprehensive study, we provide a concise guide for researchers interested in collecting or approximating human similarity data. To accompany this guide, we also release all of the similarity and language data, a total of 206,339 human judgments, that we collected in our experiments, along with a detailed breakdown of all modeling results.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext addf89a4-1826-4df2-93cc-62237bda09ecCited by top-tier papers5
- Giving Robots a Voice: Human-in-the-Loop Voice Creation and open-ended LabelingPol van Rijn, Silvan Mertes, Kathrin Janowski, Katharina Weitz et al.CHI 2024 · 12 citations
- Evaluating alignment between humans and neural network representations in image-based learning tasksCan Demircan, Tankred Saanum, Leonardo Pettini, Marcel Binz et al.NeurIPS 2024 · 11 citations
- Learning Human-like Representations to Enable Learning Human ValuesAndrea Wynn, Ilia Sucholutsky, Tom GriffithsNeurIPS 2024 · 11 citations
- Mind Your Step (by Step): Chain-of-Thought can Reduce Performance on Tasks where Thinking Makes Humans WorseRyan Liu, Jiayi Geng, Addison J. Wu, Ilia Sucholutsky et al.ICML 2025 · 4 citations
- One Hundred Neural Networks and Brains Watching Videos: Lessons from AlignmentChristina Sartzetaki, Gemma Roig, Cees G. M. Snoek, Iris I. A. GroenICLR 2025
Builds on11
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Swin Transformer: Hierarchical Vision Transformer using Shifted WindowsZe Liu, Yutong Lin, Yue Cao, Han Hu et al.ICCV 2021 · 31,683 citations
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
- Supervised Contrastive LearningPrannay Khosla, Piotr Teterwak, Chen Wang, Aaron Sarna et al.NeurIPS 2020 · 7,049 citations
Related papers
- LVM-Med: Learning Large-Scale Self-Supervised Vision Models for Medical Imaging via Second-order Graph MatchingDuy M. H. Nguyen, Hoang Nguyen, Nghiem Tuong Diep, Tan Ngoc Pham et al.NeurIPS 2023 · 107 citations
- Learning Visual Representations via Language-Guided SamplingMohamed El Banani, Karan Desai, Justin JohnsonCVPR 2023
- Integrating Language Guidance into Vision-based Deep Metric LearningKarsten Roth, Oriol Vinyals, Zeynep AkataCVPR 2022 · 1 citation
- Revealing Vision-Language Integration in the Brain with Multimodal NetworksVighnesh Subramaniam, Colin Conwell, Christopher Wang, Gabriel Kreiman et al.ICML 2024 · 19 citations
- LEMoN: Label Error Detection using Multimodal NeighborsHaoran Zhang, Aparna Balagopalan, Nassim Oufattole, Hyewon Jeong et al.ICML 2025
