Memory efficiency and resource-rational encoding in sentence processing
Weijie Xu, Brian Dillon, Richard Futrell
Abstract
There is a growing consensus that, in order to serve as models of human language processing, language models (LMs) need to be constrained in their use of memory for context, the analogue to human working memory (WM). Here we take a novel yet simple approach to constraining WM in language models, in a way that reflects models of human cognition where memory is treated as a limited resource and deployed strategically. In order to capture this constraint on memory encoding, we inject noise into the hidden representations of Transformerbased LMs at tunable rates. Then we train the models with a hybrid objective, such that they learn to maximize the performance of nextword prediction subject to explicit constraints on the total encoding precision. We find that explicit WM constraints improve the model's alignment with human reading times. More importantly, we find that the need to manage encoding precision reshapes the nature of the models' context representations, making them more compressed and categorical. Our results show how resource-rational models of WM allocation can be implemented in neural models simply and successfully, and point to a dissociation between WM retrieval mechanisms and the underlying memory representations in models of human sentence processing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 97137ae0-d664-482a-be1b-dcdd704395c8Builds on7
- QLoRA: Efficient Finetuning of Quantized LLMsTim Dettmers, Artidoro Pagnoni, Ari Holtzman, Luke ZettlemoyerNeurIPS 2023 · 5,863 citations
- Transformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space DualityTri Dao, Albert GuICML 2024 · 1,407 citations
- A Systematic Assessment of Syntactic Generalization in Neural Language ModelsJennifer Hu, Jon Gauthier, Peng Qian, Ethan Wilcox et al.ACL 2020 · 124 citations
- Context Limitations Make Neural Language Models More Human-LikeTatsuki Kuribayashi, Yohei Oseki, Ana Brassard, Kentaro InuiEMNLP 2022 · 28 citations
- Entropy- and Distance-Based Predictors From GPT-2 Attention Patterns Predict Reading Times Over and Above GPT-2 SurprisalByung-Doh Oh, William SchulerEMNLP 2022 · 13 citations
Related papers
- A Dual-Task Paradigm to Investigate Sentence Comprehension Strategies in Language ModelsRei Emura, Saku SugawaraACL 2026
- Efficient Allocation of Working Memory Resource for Utility Maximization in Humans and Recurrent Neural NetworksQingqing Yang, Hsin-Hung LiNeurIPS 2025
- Developmentally-plausible Working Memory Shapes a Critical Period for Language AcquisitionMasato Mita, Ryo Yoshida, Yohei OsekiACL 2025 · 7 citations
- Gated Differentiable Working Memory for Long-Context Language ModelingLingrui Mei, Shenghua Liu, Yiwei Wang, Yuyao Ge et al.ACL 2026 · 4 citations
- GradMem: Learning to Write Context into Memory with Test-Time Gradient DescentYuri Kuratov, Matvey Kairov, Aydar Bulatov, Ivan Rodkin et al.ICML 2026 · 3 citations
