The Visual Iconicity Challenge: Evaluating Vision-Language Models on Sign Language Form-Meaning Mapping
Onur Keles, Asli Özyürek, Gerardo Ortega, Kadir Gökgöz, Esam Ghaleb
Abstract
Iconicity, the resemblance between linguistic form and meaning, is pervasive in signed languages, offering a natural testbed for visual grounding. For vision-language models (VLMs), the challenge is to recover such essential mappings from dynamic human motion rather than static context. We introduce the Visual Iconicity Challenge, a novel video-based benchmark that adapts psycholinguistic measures to evaluate VLMs on three tasks: (i) phonological sign-form prediction (e.g., handshape, location), (ii) transparency (inferring meaning from visual form), and (iii) graded iconicity ratings. We assess 13 state-of-the-art VLMs in zero- and few-shot settings on Sign Language of the Netherlands and compare them to human baselines. On phonological form prediction, VLMs recover some handshape and location detail but remain below human performance; on transparency, they are far from human baselines; and only top models correlate moderately with human iconicity ratings. Interestingly, models with stronger phonological form prediction correlate better with human iconicity judgment, indicating shared sensitivity to visually grounded structure. Our findings validate these diagnostic tasks and motivate human-centric signals and embodied learning methods for modelling iconicity and improving visual grounding in multimodal models.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext fb4fa6ea-fea0-4bed-8e69-c9fe20c036acBuilds on7
- Kiki or Bouba? Sound Symbolism in Vision-and-Language ModelsMorris Alper, Hadar Averbuch-ElorNeurIPS 2023 · 23 citations
- VLM4D: Towards Spatiotemporal Awareness in Vision Language ModelsShijie Zhou, Alexander Vilesov, Xuehai He, Ziyu Wan et al.ICCV 2025 · 9 citations
- Learning to Generalize Without Bias for Open-Vocabulary Action RecognitionYating Yu, Congqi Cao, Yifan Zhang, Yanning ZhangICCV 2025 · 2 citations
- With Ears to See and Eyes to Hear: Sound Symbolism Experiments with Multimodal Large Language ModelsTyler Loakman, Yucheng Li, Chenghua LinEMNLP 2024 · 1 citation
- PeVL: Pose-Enhanced Vision-Language Model for Fine-Grained Human Action RecognitionHaosong Zhang, Mei Chee Leong, Liyuan Li, Weisi LinCVPR 2024
Related papers
- IGLUE: A Benchmark for Transfer Learning across Modalities, Tasks, and LanguagesEmanuele Bugliarello, Fangyu Liu, Jonas Pfeiffer, Siva Reddy et al.ICML 2022 · 71 citations
- Grounding or Guessing? Visual Signals for Detecting Hallucinations in Sign Language TranslationYasser Hamidullah, Koel Dutta Chowdhury, Yusser Al Ghussin, Shakib Yazdani et al.ICLR 2026 · 3 citations
- Reconstructing Signing Avatars from Video Using Linguistic PriorsMaria-Paola Forte, Peter Kulits, Chun-Hao Huang, Vasileios Choutas et al.CVPR 2023
- How Foundational Skills Influence VLM-based Embodied Agents: A Native PerspectiveBo Peng, Pi Bu, Keyu Pan, Xinrun Xu et al.AAAI 2026 · 1 citation
- Logos as a Well-Tempered Pre-train for Sign Language RecognitionIlya Ovodov, Petr Surovtsev, Karina Kvanchiani, Alexander Kapitanov et al.EMNLP 2025 · 1 citation
