Unveiling the mystery of visual attributes of concrete and abstract concepts: Variability, nearest neighbors, and challenging categories
Tarun Tater, Sabine Schulte im Walde, Diego Frassinelli
Abstract
The visual representation of a concept varies significantly depending on its meaning and the context where it occurs; this poses multiple challenges both for vision and multimodal models. Our study focuses on concreteness, a well-researched lexical-semantic variable, using it as a case study to examine the variability in visual representations. We rely on images associated with approximately 1,000 abstract and concrete concepts extracted from two different datasets: Bing and YFCC. Our goals are: (i) evaluate whether visual diversity in the depiction of concepts can reliably distinguish between concrete and abstract concepts; (ii) analyze the variability of visual features across multiple images of the same concept through a nearest neighbor analysis; and (iii) identify challenging factors contributing to this variability by categorizing and annotating images. Our findings indicate that for classifying images of abstract versus concrete concepts, a combination of basic visual features such as color and texture is more effective than features extracted by more complex models like Vision Transformer (ViT). However, ViTs show better performances in the nearest neighbor analysis, emphasizing the need for a careful selection of visual features when analyzing conceptual variables through modalities other than text.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 6db20ef9-3867-49ab-942a-a22ac4c15f6fCited by top-tier papers2
- AbsVis - Benchmarking How Humans and Vision-Language Models "See" Abstract Concepts in ImagesTarun Tater, Diego Frassinelli, Sabine Schulte im WaldeEMNLP 2025 · 2 citations
- Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding SpaceSi Wu, Sebastian BruchACL 2025
Builds on3
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- GeneCIS: A Benchmark for General Conditional Image SimilaritySagar Vaze, Nicolas Carion, Ishan MisraCVPR 2023
Related papers
- Seeing the Abstract: Translating the Abstract Language for Vision Language ModelsDavide Talon, Federico Girella, Ziyue Liu, Marco Cristani et al.CVPR 2025
- Exploring Concreteness Through a Figurative LensSaptarshi Ghosh, Tianyu JiangACL 2026
- ARC Is a Vision Problem!Keya Hu, Ali Cy, Linlu Qiu, Xiaoman Delores Ding et al.CVPR 2026 · 23 citations
- Beyond the Doors of Perception: Vision Transformers Represent Relations Between ObjectsMichael A. Lepori, Alexa R. Tartaglini, Wai Keen Vong, Thomas Serre et al.NeurIPS 2024 · 22 citations
- How can embedding models bind concepts?Arnas Uselis, Darina Koishigarina, Seong Joon OhICML 2026
