Where the Cat Sat: A Multilingual Framework for Spatial Language Understanding
Demian Inostroza Améstica, Ekaterina Vylomova, Charles Kemp, Mae Carroll, Wanchun Li, Meladel Mistica
摘要
Spatial language understanding is fundamental to tasks from navigation and direction, robot control, to document understanding for a multitude of languages and tasks, yet current work exhibits biases toward English and prepositional marking. We present a multilingual framework and benchmark decomposing spatial relations into surface elements (figure, ground, predicate, markers) and semantic components (dynamicity, stasis). Evaluating frontier large language models (LLMs) on Spanish, Basque, and Chinese with text-only input, we find high performance on figure and ground identification but persistent gaps in two areas: semantic classification of topological and projective relations, and surface identification of morphological spatial markers-Basque case affixes proving most challenging with recognition of spatial elements as low as 15.3%. These results suggest that surface parsing does not entail spatial understanding, and that evaluation must include spatial marking strategies used across typologically diverse languages.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper3
- Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMsShengbang Tong, Zhuang Liu, Yuexiang Zhai, Yi Ma 等CVPR 2024 · 被引用 111 次
- Can Multimodal Large Language Models Understand Spatial Relations?Jingping Liu, Ziyan Liu, Zhedong Cen, Yan Zhou 等ACL 2025 · 被引用 16 次
- Do Vision-Language Models Represent Space and How? Evaluating Spatial Frame of Reference under AmbiguitiesZheyuan Zhang, Fengyuan Hu, Jayjun Lee, Freda Shi 等ICLR 2025
相关 Paper
- STI-Bench: Are MLLMs Ready for Precise Spatial-Temporal World Understanding?Yun Li, Yiming Zhang, Tao Lin, XiangRui Liu 等ICCV 2025 · 被引用 4 次
- The Same but Different: Structural Similarities and Differences in Multilingual Language ModelingRuochen Zhang, Qinan Yu, Matianyu Zang, Carsten Eickhoff 等ICLR 2025
- Do LLMs Capture Embodied Cognition and Cultural Variation? Cross-Linguistic Evidence from DemonstrativesYu Wang, Emmanuele Chersoni, Chu-Ren HuangACL 2026
- SpatialLM: Training Large Language Models for Structured Indoor ModelingYongsen Mao, Junhao Zhong, Chuan Fang, Jia Zheng 等NeurIPS 2025 · 被引用 89 次
- Multi-SpatialMLLM: Multi-Frame Spatial Understanding with Multi-Modal Large Language ModelsRunsen Xu, Weiyao Wang, Hao Tang, Xingyu Chen 等CVPR 2026 · 被引用 64 次
