Uncovering Visual-Semantic Psycholinguistic Properties from the Distributional Structure of Text Embedding Space
Si Wu, Sebastian Bruch
摘要
Imageability (potential of text to evoke a mental image) and concreteness (perceptibility of text) are two psycholinguistic properties that link visual and semantic spaces. It is little surprise that computational methods that estimate them do so using parallel visual and semantic spaces, such as collections of image-caption pairs or multi-modal models. In this paper, we work on the supposition that text itself in an image-caption dataset offers sufficient signals to accurately estimate these properties. We hypothesize, in particular, that the peakedness of the neighborhood of a word in the semantic embedding space reflects its degree of imageability and concreteness. We then propose an unsupervised, distribution-free measure, which we call Neighborhood Stability Measure (NSM), that quantifies the sharpness of peaks. Extensive experiments show that NSM correlates more strongly with ground-truth ratings than existing unsupervised methods, and is a strong predictor of these properties for classification. Our code and data are available on GitHub (https://github.com/Artificial-Memory-Lab/imageability).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- Unveiling the mystery of visual attributes of concrete and abstract concepts: Variability, nearest neighbors, and challenging categoriesTarun Tater, Sabine Schulte im Walde, Diego FrassinelliEMNLP 2024 · 被引用 2 次
- Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual ConceptsSoravit Changpinyo, Piyush Sharma, Nan Ding, Radu SoricutCVPR 2021
相关 Paper
- Investigating Conceptual Blending of a Diffusion Model for Improving Nonword-to-Image GenerationChihaya Matsuhira, Marc A. Kastner, Takahiro Komamizu, Takatsugu Hirayama 等ACM MM 2024 · 被引用 1 次
- Word-As-Image for Semantic TypographyShir Iluz, Yael Vinker, Amir Hertz, Daniel Berio 等SIGGRAPH 2023 · 被引用 67 次
- Exploring Concreteness Through a Figurative LensSaptarshi Ghosh, Tianyu JiangACL 2026
- Urban2Vec: Incorporating Street View Imagery and POIs for Multi-Modal Urban Neighborhood EmbeddingZhecheng Wang, Haoyuan Li, Ram RajagopalAAAI 2020 · 被引用 113 次
- Quantifying Learnability and Describability of Visual Concepts Emerging in Representation LearningIro Laina, Ruth Fong, Andrea VedaldiNeurIPS 2020 · 被引用 15 次
