With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
Tyler Loakman, Yucheng Li, Chenghua Lin
摘要
Recently, Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated aptitude as potential substitutes for human participants in experiments testing psycholinguistic phenomena. However, an understudied question is to what extent models that only have access to vision and text modalities are able to implicitly understand sound-based phenomena via abstract reasoning from orthography and imagery alone. To investigate this, we analyse the ability of VLMs and LLMs to demonstrate sound symbolism (i.e., to recognise a non-arbitrary link between sounds and concepts) as well as their ability to “hear” via the interplay of the language and vision modules of open and closed-source multimodal models. We perform multiple experiments, including replicating the classic Kiki-Bouba and Mil-Mal shape and magnitude symbolism tasks and comparing human judgements of linguistic iconicity with that of LLMs. Our results show that VLMs demonstrate varying levels of agreement with human labels, and more task information may be required for VLMs versus their human counterparts for in silico experimentation. We additionally see through higher maximum agreement levels that Magnitude Symbolism is an easier pattern for VLMs to identify than Shape Symbolism, and that an understanding of linguistic iconicity is highly dependent on model size.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki EffectTom Kouwenhoven, Kiana Shahrasbi, Tessa VerhoefNeurIPS 2025 · 被引用 4 次
- Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound SymbolismJinhong Jeong, Sunghyun Lee, Jaeyoung Lee, Seonah Han 等AAAI 2026
- The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning MappingOnur Keles, Asli Özyürek, Gerardo Ortega, Kadir Gökgöz 等ACL 2026
它引用的顶会 Paper4
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 被引用 651 次
- Knowledge of cultural moral norms in large language modelsAida Ramezani, Yang XuACL 2023 · 被引用 44 次
- Kiki or Bouba? Sound Symbolism in Vision-and-Language ModelsMorris Alper, Hadar Averbuch-ElorNeurIPS 2023 · 被引用 23 次
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
相关 Paper
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 被引用 1 次
- Can Large Language Models Understand Symbolic Graphics Programs?Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu 等ICLR 2025
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen 等ICLR 2026 · 被引用 103 次
- Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard ProblemsMikolaj Malkinski, Szymon Pawlonka, Jacek MandziukICML 2025
- Language Models Don't Learn the Physical Manifestation of LanguageBruce W. Lee, Jaehyuk LimACL 2024
