Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space
Si Wu, Sebastian Bruch
Abstract
Imageability (potential of text to evoke a mental image) and concreteness (perceptibility of text) are two psycholinguistic properties that link visual and semantic spaces. It is little surprise that computational methods that estimate them do so using parallel visual and semantic spaces, such as collections of image-caption pairs or multi-modal models. In this paper, we work on the supposition that text itself in an image-caption dataset offers sufficient signals to accurately estimate these properties. We hypothesize, in particular, that the peakedness of the neighborhood of a word in the semantic embedding space reflects its degree of imageability and concreteness. We then propose an unsupervised, distribution-free measure, which we call Neighborhood Stability Measure (NSM), that quantifies the sharpness of peaks. Extensive experiments show that NSM correlates more strongly with ground-truth ratings than existing unsupervised methods, and is a strong predictor of these properties for classification. Our code and data are available on GitHub (https://github.com/Artificial-Memory-Lab/imageability).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on3
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- Unveiling the mystery of visual attributes of concrete and abstract concepts: Variability, nearest neighbors, and challenging categoriesTarun Tater, Sabine Schulte im Walde, Diego FrassinelliEMNLP 2024 · 2 citations
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual ConceptsSoravit Changpinyo, Piyush Sharma, Nan Ding, Radu SoricutCVPR 2021
Related papers
- Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image GenerationChihaya Matsuhira, Marc A. Kastner, Takahiro Komamizu, Takatsugu Hirayama et al.ACM MM 2024 · 1 citation
- Word-As-Image for Semantic TypographyShir Iluz, Yael Vinker, Amir Hertz, Daniel Berio et al.SIGGRAPH 2023 · 67 citations
- Exploring Concreteness Through a Figurative LensSaptarshi Ghosh, Tianyu JiangACL 2026
- Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood EmbeddingZhecheng Wang, Haoyuan Li, Ram RajagopalAAAI 2020 · 113 citations
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation LearningIro Laina, Ruth Fong, Andrea VedaldiNeurIPS 2020 · 15 citations
