With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language Models
Tyler Loakman, Yucheng Li, Chenghua Lin
Abstract
Recently, Large Language Models (LLMs) and Vision Language Models (VLMs) have demonstrated aptitude as potential substitutes for human participants in experiments testing psycholinguistic phenomena. However, an understudied question is to what extent models that only have access to vision and text modalities are able to implicitly understand sound-based phenomena via abstract reasoning from orthography and imagery alone. To investigate this, we analyse the ability of VLMs and LLMs to demonstrate sound symbolism (i.e., to recognise a non-arbitrary link between sounds and concepts) as well as their ability to “hear” via the interplay of the language and vision modules of open and closed-source multimodal models. We perform multiple experiments, including replicating the classic Kiki-Bouba and Mil-Mal shape and magnitude symbolism tasks and comparing human judgements of linguistic iconicity with that of LLMs. Our results show that VLMs demonstrate varying levels of agreement with human labels, and more task information may be required for VLMs versus their human counterparts for in silico experimentation. We additionally see through higher maximum agreement levels that Magnitude Symbolism is an easier pattern for VLMs to identify than Shape Symbolism, and that an understanding of linguistic iconicity is highly dependent on model size.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext db408bab-baf1-446f-8d5a-b0744ed4b7b0Cited by top-tier papers3
- Cross-modal Associations in Vision and Language Models: Revisiting the Bouba-Kiki EffectTom Kouwenhoven, Kiana Shahrasbi, Tessa VerhoefNeurIPS 2025 · 4 citations
- Do Language Models Associate Sound with Meaning? A Multimodal Study of Sound SymbolismJinhong Jeong, Sunghyun Lee, Jaeyoung Lee, Seonah Han et al.AAAI 2026
- The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning MappingOnur Keles, Asli Özyürek, Gerardo Ortega, Kadir Gökgöz et al.ACL 2026
Builds on4
- Using Large Language Models to Simulate Multiple Humans and Replicate Human Subject StudiesGati V. Aher, Rosa I. Arriaga, Adam Tauman KalaiICML 2023 · 651 citations
- Knowledge of cultural moral norms in large language modelsAida Ramezani, Yang XuACL 2023 · 44 citations
- Kiki or Bouba? Sound Symbolism in Vision-and-Language ModelsMorris Alper, Hadar Averbuch-ElorNeurIPS 2023 · 23 citations
- Improved Baselines with Visual Instruction TuningHaotian Liu, Chunyuan Li, Yuheng Li, Yong Jae LeeCVPR 2024
Related papers
- VisNumBench: Evaluating Number Sense of Multimodal Large Language ModelsTengjin Weng, Jingyi Wang, Wenhao Jiang, Zhong MingICCV 2025 · 1 citation
- Can Large Language Models Understand Symbolic Graphics Programs?Zeju Qiu, Weiyang Liu, Haiwen Feng, Zhen Liu et al.ICLR 2025
- VisuLogic: A Benchmark for Evaluating Visual Reasoning in Multi-modal Large Language ModelsWeiye Xu, Jiahao Wang, Weiyun Wang, Zhe Chen et al.ICLR 2026 · 103 citations
- Reasoning Limitations of Multimodal Large Language Models. A case study of Bongard ProblemsMikolaj Malkinski, Szymon Pawlonka, Jacek MandziukICML 2025
- Language Models Don't Learn the Physical Manifestation of LanguageBruce W. Lee, Jaehyuk LimACL 2024
