Interpretable Measures of Conceptual Similarity by Complexity-Constrained Descriptive Auto-Encoding
Alessandro Achille, Greg Ver Steeg, Tian Yu Liu, Matthew Trager, Carson Klingenberg, Stefano Soatto
Abstract
Quantifying the degree of similarity between images is a key copyright issue for image-based machine learning. In legal doctrine however, determining the degree of similarity between works requires subjective analysis, and fact-finders (judges and juries) can demonstrate considerable variability in these subjective judgement calls. Images that are structurally similar can be deemed dissimilar, whereas images of completely different scenes can be deemed similar enough to support a claim of copying. We seek to define and compute a notion of 'conceptual similarity' among images that captures high-level relations even among images that do not share repeated elements or visually similar components. The idea is to use a base multi-modal model to generate 'explanations' (captions) of visual data at increasing levels of complexity. Then, similarity can be measured by the length of the caption needed to discriminate between the two images: Two highly dissimilar images can be discriminated early in their description, whereas conceptually dissimilar ones will need more detail to be distinguished. We operationalize this definition and show that it correlates with subjective (averaged human evaluation) assessment, and beats existing baselines on both image-to-image and text-to-text similarity benchmarks. Beyond just providing a number, our method also offers interpretability by pointing to the specific level of granularity of the description where the source data are differentiated.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 30a24f2c-1c23-4f89-bf59-76ea0b870fddCited by top-tier papers1
Ask how each one uses itBuilds on10
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Visual Instruction TuningHaotian Liu, Chunyuan Li, Qingyang Wu, Yong Jae LeeNeurIPS 2023 · 11,349 citations
- Concept Bottleneck ModelsPang Wei Koh, Thao Nguyen, Yew Siang Tang, Stephen Mussmann et al.ICML 2020 · 1,233 citations
- PromptBERT: Improving BERT Sentence Embeddings with PromptsTing Jiang, Jian Jiao, Shaohan Huang, Zihan Zhang et al.EMNLP 2022 · 148 citations
- An Unsupervised Sentence Embedding Method by Mutual Information MaximizationYan Zhang, Ruidan He, Zuozhu Liu, Kwan Hui Lim et al.EMNLP 2020 · 126 citations
Related papers
- Structure Your Data: Towards Semantic Graph CounterfactualsAngeliki Dimitriou, Maria Lymperaiou, Giorgos Filandrianos, Konstantinos Thomas et al.ICML 2024 · 7 citations
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation LearningIro Laina, Ruth Fong, Andrea VedaldiNeurIPS 2020 · 15 citations
- Relational Visual SimilarityThao Nguyen, Sicheng Mo, Krishna Kumar Singh, Yilin Wang et al.CVPR 2026 · 1 citation
- DreamSim: Learning New Dimensions of Human Visual Similarity using Synthetic DataStephanie Fu, Netanel Tamir, Shobhita Sundaram, Lucy Chai et al.NeurIPS 2023 · 413 citations
- Representational Similarity via Interpretable Visual ConceptsNeehar Kondapaneni, Oisin Mac Aodha, Pietro PeronaICLR 2025
