Using Perspectival Words Is Harder Than Vocabulary Words for Humans - and Even More So for Multimodal Language Models
Dota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Özyürek, Paula Rubio-Fernández
Abstract
Multimodal language models (MLMs) increasingly demonstrate human-like communication, yet their use of everyday perspectival words remains poorly understood. To address this gap, we compare humans and MLMs in their use of three word types, which we predict impose increasing cognitive demands: vocabulary (e.g., 'boat' or 'cup'), possessives (e.g., 'mine' vs. 'yours'), and demonstratives (e.g., 'this one' vs. 'that one'). Testing seven MLMs against human participants, we find that perspectival words are harder than vocabulary words for both groups. The gap is even larger for MLMs: while models approach human-level performance on using vocabulary, they exhibit clear deficits with possessives and even greater difficulties with demonstratives. Ablation analyses point to limitations in perspective-taking and spatial reasoning as key sources of these gaps in MLMs. Instruction-based prompting helps close the gap for possessives but still leaves demonstratives far below human performance. These results show that, unlike vocabulary, perspectival words pose a greater challenge in human communication-and this difficulty is further amplified in MLMs, revealing a crucial shortfall in their pragmatic and social-cognitive abilities. * Equal contribution † Work has been done as an student assistant at MPI Codes available at https://github.com/Beckinetic/ VLMIndexical * 'Perspectival words' is a non-technical label for indexicals, specifically personal, possessive, and demonstrative pronouns.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on9
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- Tree of Thoughts: Deliberate Problem Solving with Large Language ModelsShunyu Yao, Dian Yu, Jeffrey Zhao, Izhak Shafran et al.NeurIPS 2023 · 5,068 citations
- Demystifying CLIP DataHu Xu, Saining Xie, Xiaoqing Ellen Tan, Po-Yao Huang et al.ICLR 2024 · 249 citations
- IMPLI: Investigating NLI Models' Performance on Figurative LanguageKevin Stowe, Prasetya Ajie Utama, Iryna GurevychACL 2022 · 52 citations
- A fine-grained comparison of pragmatic language understanding in humans and language modelsJennifer Hu, Sammy Floyd, Olessia Jouravlev, Evelina Fedorenko et al.ACL 2023 · 45 citations
Related papers
- Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from DemonstrativesYu Wang, Emmanuele Chersoni, Chu-Ren HuangACL 2026
- SPHERE: Unveiling Spatial Blind Spots in Vision-Language Models Through Hierarchical EvaluationWenyu Zhang, Wei En Ng, Lixin Ma, Yuwen Wang et al.ACL 2025 · 20 citations
- EgoProx: Evaluating MLLMs on Egocentric 3D Proximity Reasoning Across a Cognitive HierarchyJinzhao Li, Yinuo Chen, Dongxu Piao, Panwang Pan et al.CVPR 2026 · 2 citations
- TopViewRS: Vision-Language Models as Top-View Spatial ReasonersChengzu Li, Caiqi Zhang, Han Zhou, Nigel Collier et al.EMNLP 2024 · 5 citations
- Spatial CAPTCHA: Generatively Benchmarking Spatial Reasoning for Human-Machine DifferentiationArina Kharlamova, Bowei He, Chen Ma, Xue LiuICLR 2026 · 3 citations
