Dealing with Semantic Underspecification in Multimodal NLP
Sandro Pezzelle
Abstract
Intelligent systems that aim at mastering language as humans do must deal with its semantic underspecification, namely, the possibility for a linguistic signal to convey only part of the information needed for communication to succeed. Consider the usages of the pronoun they, which can leave the gender and number of its referent(s) underspecified. Semantic underspecification is not a bug but a crucial language feature that boosts its storage and processing efficiency. Indeed, human speakers can quickly and effortlessly integrate semantically-underspecified linguistic signals with a wide range of non-linguistic information, e.g., the multimodal context, social or cultural conventions, and shared knowledge. Standard NLP models have, in principle, no or limited access to such extra information, while multimodal systems grounding language into other modalities, such as vision, are naturally equipped to account for this phenomenon. However, we show that they struggle with it, which could negatively affect their performance and lead to harmful consequences when used for applications. In this position paper, we argue that our community should be aware of semantic underspecification if it aims to develop language technology that can successfully interact with human users. We discuss some applications where mastering it is crucial and outline a few directions toward achieving this goal.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 07e27ce8-a3b4-4049-b1ad-bedcb07579c0Cited by top-tier papers9
- LLMs Get Lost In Multi-Turn ConversationPhilippe Laban, Hiroaki Hayashi, Yingbo Zhou, Jennifer NevilleICLR 2026 · 491 citations
- Emergence of Hidden Capabilities: Exploring Learning Dynamics in Concept SpaceCore Francisco Park, Maya Okawa, Andrew Lee, Ekdeep Singh Lubana et al.NeurIPS 2024 · 39 citations
- We're Afraid Language Models Aren't Modeling AmbiguityAlisa Liu, Zhaofeng Wu, Julian Michael, Alane Suhr et al.EMNLP 2023 · 35 citations
- Rephrase, Augment, Reason: Visual Grounding of Questions for Vision-Language ModelsArchiki Prasad, Elias Stengel-Eskin, Mohit BansalICLR 2024 · 13 citations
- ContextRef: Evaluating Referenceless Metrics for Image Description GenerationElisa Kreiss, Eric Zelikman, Christopher Potts, Nick HaberICLR 2024 · 6 citations
Builds on20
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- ViLT: Vision-and-Language Transformer Without Convolution or Region SupervisionWonjae Kim, Bokyung Son, Ildoo KimICML 2021 · 2,258 citations
- VL-BERT: Pre-training of Generic Visual-Linguistic RepresentationsWeijie Su, Xizhou Zhu, Yue Cao, Bin Li et al.ICLR 2020 · 1,825 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
Related papers
- Correct-Detect: Balancing Performance and Ambiguity Through the Lens of Coreference Resolution in LLMsAmber Shore, Russell Scheinberg, Ameeta Agrawal, So Young LeeEMNLP 2025
- Using Perspectival Words Is Harder Than Vocabulary Words for Humans - and Even More So for Multimodal Language ModelsDota Tianai Dong, Yifan Luo, Po-Ya Angela Wang, Asli Özyürek et al.ACL 2026
- Advancing Social Intelligence in AI Agents: Technical Challenges and Open QuestionsLeena Mathur, Paul Pu Liang, Louis-Philippe MorencyEMNLP 2024 · 6 citations
- Plug-and-Play Clarifier: A Zero-Shot Multimodal Framework for Egocentric Intent DisambiguationSicheng Yang, Yukai Huang, Weitong Cai, Shitong Sun et al.AAAI 2026
- VAGUE: Visual Contexts Clarify Ambiguous ExpressionsHeejeong Nam, Jinwoo Ahn, Keummin Ka, Jiwan Chung et al.ICCV 2025 · 1 citation
