Mapping Language Models to Grounded Conceptual Spaces
Roma Patel, Ellie Pavlick
摘要
A fundamental criticism of text-only language models (LMs) is their lack of grounding---that is, the ability to tie a word for which they have learned a representation, to its actual use in the world. However, despite this limitation, large pre-trained LMs have been shown to have a remarkable grasp of the conceptual structure of language, as demonstrated by their ability to answer questions, generate fluent text, or make inferences about entities, objects, and properties that they have never physically observed. In this work we investigate the extent to which the rich conceptual structure that LMs learn indeed reflects the conceptual structure of the non-linguistic world---which is something that LMs have never observed. We do this by testing whether the LMs can learn to map an entire conceptual domain (e.g., direction or colour) onto a grounded world representation given only a small number of examples. For example, we show a model what the word left" means using a textual depiction of a grid world, and assess how well it can generalise to related concepts, for example, the wordright", in a similar grid world. We investigate a range of generative language models of varying sizes (including GPT-2 and GPT-3), and see that although the smaller models struggle to perform this mapping, the largest model can not only learn to ground the concepts that it is explicitly taught, but appears to generalise to several instances of unseen concepts as well. Our results suggest an alternative means of building grounded language models: rather than learning grounded representations ``from scratch'', it is possible that large text-only models learn a sufficiently rich conceptual structure that could allow them to be grounded in a data-efficient way.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper52
- Can Language Models Solve Graph Problems in Natural Language?Heng Wang, Shangbin Feng, Tianxing He, Zhaoxuan Tan 等NeurIPS 2023 · 被引用 420 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- The Alignment Problem from a Deep Learning PerspectiveRichard Ngo, Lawrence Chan, Sören MindermannICLR 2024 · 被引用 296 次
- Grounding Large Language Models in Interactive Environments with Online Reinforcement LearningThomas Carta, Clément Romac, Thomas Wolf, Sylvain Lamprier 等ICML 2023 · 被引用 258 次
- Patchscopes: A Unifying Framework for Inspecting Hidden Representations of Language ModelsAsma Ghandeharioun, Avi Caciularu, Adam Pearce, Lucas Dixon 等ICML 2024 · 被引用 197 次
相关 Paper
- Evaluating the Effectiveness of Large Language Models in Establishing Conversational GroundingBiswesh Mohapatra, Manav Nitin Kapadnis, Laurent Romary, Justine CassellEMNLP 2024 · 被引用 1 次
- World-to-Words: Grounded Open Vocabulary Acquisition through Fast Mapping in Vision-Language ModelsZiqiao Ma, Jiayi Pan, Joyce ChaiACL 2023 · 被引用 4 次
- Conceptual structure coheres in human cognition but not in large language modelsSiddharth Suresh, Kushin Mukherjee, Xizheng Yu, Wei-Chun Huang 等EMNLP 2023 · 被引用 7 次
- A Vision Check-up for Language ModelsPratyusha Sharma, Tamar Rott Shaham, Manel Baradad, Adrián Rodríguez-Muñoz 等CVPR 2024 · 被引用 10 次
- COPEN: Probing Conceptual Knowledge in Pre-trained Language ModelsHao Peng, Xiaozhi Wang, Shengding Hu, Hailong Jin 等EMNLP 2022 · 被引用 16 次
