Memory efficiency and resource-rational encoding in sentence processing
Weijie Xu, Brian Dillon, Richard Futrell
摘要
There is a growing consensus that, in order to serve as models of human language processing, language models (LMs) need to be constrained in their use of memory for context, the analogue to human working memory (WM). Here we take a novel yet simple approach to constraining WM in language models, in a way that reflects models of human cognition where memory is treated as a limited resource and deployed strategically. In order to capture this constraint on memory encoding, we inject noise into the hidden representations of Transformerbased LMs at tunable rates. Then we train the models with a hybrid objective, such that they learn to maximize the performance of nextword prediction subject to explicit constraints on the total encoding precision. We find that explicit WM constraints improve the model's alignment with human reading times. More importantly, we find that the need to manage encoding precision reshapes the nature of the models' context representations, making them more compressed and categorical. Our results show how resource-rational models of WM allocation can be implemented in neural models simply and successfully, and point to a dissociation between WM retrieval mechanisms and the underlying memory representations in models of human sentence processing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 被引用 5,863 次
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 被引用 1,407 次
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox 等ACL 2020 · 被引用 124 次
- Context Limitations Make Neural Language Models More Human-LikeTatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro InuiEMNLP 2022 · 被引用 28 次
- Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 SurprisalByung-Doh Oh, William SchulerEMNLP 2022 · 被引用 13 次
相关 Paper
- A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language ModelsRei Emura, Saku SugawaraACL 2026
- Efficient Allocation of Working Memory Resource for Utility Maximization in Humans and Recurrent Neural NetworksQingqing Yang, Hsin-Hung LiNeurIPS 2025
- Developmentally-plausible Working Memory Shapes a Critical Period for Language AcquisitionMasato Mita, Ryo Yoshida, Yohei OsekiACL 2025 · 被引用 7 次
- Gated Differentiable Working Memory for Long-Context Language ModelingLingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge 等ACL 2026 · 被引用 4 次
- GradMem: Learning to Write Context into Memory with Test-Time Gradient DescentYuri Kuratov, Matvey Kairov, Aydar Bulatov, Ivan Rodkin 等ICML 2026 · 被引用 3 次
