Describe Me an Auklet: Generating Grounded Perceptual Category Descriptions
Bill Noble, Nikolai Ilinykh
Abstract
Human speakers can generate descriptions of perceptual concepts, abstracted from the instance-level. Moreover, such descriptions can be used by other speakers to learn provisional representations of those concepts. Learning and using abstract perceptual concepts is under-investigated in the language-and-vision field. The problem is also highly relevant to the field of representation learning in multi-modal NLP. In this paper, we introduce a framework for testing category-level perceptual grounding in multi-modal language models. In particular, we train separate neural networks to generate and interpret descriptions of visual categories. We measure the communicative success of the two models with the zero-shot classification performance of the interpretation model, which we argue is an indicator of perceptual grounding. Using this framework, we compare the performance of prototype- and exemplar-based representations. Finally, we show that communicative success exposes performance issues in the generation model, not captured by traditional intrinsic NLG evaluation metrics, and argue that these issues stem from a failure to properly ground language in vision at the category level.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on6
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Mapping Language Models to Grounded Conceptual SpacesRoma Patel, Ellie PavlickICLR 2022 · 197 citations
- Experience Grounds LanguageYonatan Bisk, Ari Holtzman, Jesse Thomason, Jacob Andreas et al.EMNLP 2020 · 74 citations
Related papers
- CompGuessWhat?!: A Multi-task Evaluation Framework for Grounded Language LearningAlessandro Suglia, Ioannis Konstas, Andrea Vanzo, Emanuele Bastianelli et al.ACL 2020 · 12 citations
- Seeing Beyond Classes: Zero-Shot Grounded Situation Recognition via Language ExplainerJiaming Lei, Lin Li, Chunping Wang, Jun Xiao et al.ACM MM 2024
- Waffling around for Performance: Visual Classification with Random Words and Broad ConceptsKarsten Roth, Jae-Myung Kim, A. Sophia Koepke, Oriol Vinyals et al.ICCV 2023 · 124 citations
- Learning Cross-Modal Context Graph for Visual GroundingYongfei Liu, Bo Wan, Xiaodan Zhu, Xuming HeAAAI 2020 · 100 citations
- X-SAM: From Segment Anything to Any SegmentationHao Wang, Limeng Qiao, Zequn Jie, Zhijian Huang et al.AAAI 2026 · 16 citations
