Kiki or Bouba? Sound Symbolism in Vision-and-Language Models
Morris Alper, Hadar Averbuch-Elor
摘要
Although the mapping between sound and meaning in human language is assumed to be largely arbitrary, research in cognitive science has shown that there are non-trivial correlations between particular sounds and meanings across languages and demographic groups, a phenomenon known as sound symbolism. Among the many dimensions of meaning, sound symbolism is particularly salient and welldemonstrated with regards to cross-modal associations between language and the visual domain. In this work, we address the question of whether sound symbolism is reflected in vision-and-language models such as CLIP and Stable Diffusion. Using zero-shot knowledge probing to investigate the inherent knowledge of these models, we find strong evidence that they do show this pattern, paralleling the well-known kiki-bouba effect in psycholinguistics. Our work provides a novel method for demonstrating sound symbolism and understanding its nature using computational tools. Our code will be made publicly available 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki EffectTom Kouwenhoven, Kiana Shahrasbi, Tessa VerhoefNeurIPS 2025 · 被引用 4 次
- ConlangCrafter: Constructing Languages with a Multi-Hop LLM PipelineMorris Alper, Moran Yanuka, Raja Giryes, Gasper BegusACL 2026 · 被引用 4 次
- Verbalized Representation Learning for Interpretable Few-Shot GeneralizationCheng-Fu Yang, Da Yin, Wenbo Hu, Heng Ji 等ICCV 2025 · 被引用 1 次
- How Tokenization Limits Phonological Knowledge Representation in Language Models and How to Improve ThemDisen Liao, Freda ShiACL 2026 · 被引用 1 次
- With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language ModelsTyler Loakman, Yucheng Li, Chenghua LinEMNLP 2024 · 被引用 1 次
它引用的顶会 Paper13
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh 等ICML 2021 · 被引用 47,906 次
- High-Resolution Image Synthesis with Latent Diffusion ModelsRobin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser 等CVPR 2022 · 被引用 13,123 次
- Photorealistic Text-to-Image Diffusion Models with Deep Language UnderstandingChitwan Saharia, William Chan, Saurabh Saxena, Lala Li 等NeurIPS 2022 · 被引用 8,965 次
- ALBERT: A Lite BERT for Self-supervised Learning of Language RepresentationsZhenzhong Lan, Mingda Chen, Sebastian Goodman, Kevin Gimpel 等ICLR 2020 · 被引用 7,418 次
- MiniLM: Deep Self-Attention Distillation for Task-Agnostic Compression of Pre-Trained TransformersWenhui Wang, Furu Wei, Li Dong, Hangbo Bao 等NeurIPS 2020 · 被引用 2,727 次
相关 Paper
- Text-to-Image Diffusion Models are Zero Shot ClassifiersKevin Clark, Priyank JainiNeurIPS 2023 · 被引用 192 次
- CLIPSep: Learning Text-queried Sound Separation with Noisy Unlabeled VideosHao-Wen Dong, Naoya Takahashi, Yuki Mitsufuji, Julian J. McAuley 等ICLR 2023 · 被引用 3 次
- Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound SymbolismJinhong Jeong, Sunghyun Lee, Jaeyoung Lee, Seonah Han 等AAAI 2026
- Images that Sound: Composing Images and Sounds on a Single CanvasZiyang Chen, Daniel Geng, Andrew OwensNeurIPS 2024 · 被引用 22 次
- PolCLIP: A Unified Image-Text Word Sense Disambiguation Model via Generating Multimodal Complementary RepresentationsQihao Yang, Yong Li, Xuelin Wang, Fu Lee Wang 等ACL 2024
