Implicit Representations of Meaning in Neural Language Models
Belinda Z. Li, Maxwell I. Nye, Jacob Andreas
Abstract
Does the effectiveness of neural language models derive entirely from accurate modeling of surface word co-occurrence statistics, or do these models represent and reason about the world they describe? In BART and T5 transformer language models, we identify contextual word representations that function as models of entities and situations as they evolve throughout a discourse. These neural representations have functional similarities to linguistic models of dynamic semantics: they support a linear readout of each entity's current properties and relations, and can be manipulated with predictable effects on language generation. Our results indicate that prediction in pretrained neural language models is supported, at least in part, by dynamic representations of meaning and implicit simulation of entity state, and that this behavior can be learned with only text as training data. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 435d71cc-39c1-4e88-975d-266fec0eeb1eCited by top-tier papers68
- Implicit In-context LearningZhuowei Li, Zihao Xu, Ligong Han, Yunhe Gao et al.ICLR 2025 · 1,989 citations
- Language Models as Zero-Shot Planners: Extracting Actionable Knowledge for Embodied AgentsWenlong Huang, Pieter Abbeel, Deepak Pathak, Igor MordatchICML 2022 · 1,539 citations
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka et al.NeurIPS 2024 · 1,166 citations
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 303 citations
- Physics of Language Models: Part 3.1, Knowledge Storage and ExtractionZeyuan Allen-Zhu, Yuanzhi LiICML 2024 · 258 citations
Builds on3
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad et al.ACL 2020 · 1,224 citations
- Climbing towards NLU: On Meaning, Form, and Understanding in the Age of DataEmily M. Bender, Alexander KollerACL 2020 · 914 citations
- Pareto Probing: Trading Off Accuracy for ComplexityTiago Pimentel, Naomi Saphra, Adina Williams, Ryan CotterellEMNLP 2020 · 6 citations
Related papers
- Entity Tracking in Language ModelsNajoung Kim, Sebastian SchusterACL 2023 · 9 citations
- Generating Coherent Narratives by Learning Dynamic and Discrete Entity States with a Contrastive FrameworkJian Guan, Zhenyu Yang, Rongsheng Zhang, Zhipeng Hu et al.AAAI 2023 · 11 citations
- Dynamic Contextualized Word EmbeddingsValentin Hofmann, Janet B. Pierrehumbert, Hinrich SchützeACL 2021
- A Model of the Language ProcessBrandon Duderstadt, Hayden S. HelmACL 2026 · 1 citation
- Long Text Generation by Modeling Sentence-Level and Discourse-Level CoherenceJian Guan, Xiaoxi Mao, Changjie Fan, Zitao Liu et al.ACL 2021
