Developmentally-plausible Working Memory Shapes a Critical Period for Language Acquisition
Masato Mita, Ryo Yoshida, Yohei Oseki
Abstract
Large language models possess general linguistic abilities but acquire language less efficiently than humans. This study proposes a method for integrating the developmental characteristics of working memory during the critical period, a stage when human language acquisition is particularly efficient, into the training process of language models. The proposed method introduces a mechanism that initially constrains working memory during the early stages of training and gradually relaxes this constraint in an exponential manner as learning progresses. Targeted syntactic evaluation shows that the proposed method outperforms conventional methods without memory constraints or with static memory constraints. These findings not only provide new directions for designing data-efficient language models but also offer indirect evidence supporting the role of the developmental characteristics of working memory as the underlying mechanism of the critical period in language acquisition.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d31c7eb1-ee0c-4633-9776-183f953b0bbdCited by top-tier papers2
- CxMP: A Linguistic Minimal-Pair Benchmark for Evaluating Constructional Understanding in Language ModelsMiyu Oba, Saku SugawaraACL 2026
- Memory efficiency and resource-rational encoding in sentence processingWeijie Xu, Brian Dillon, Richard FutrellACL 2026
Builds on3
- Train Short, Test Long: Attention with Linear Biases Enables Input Length ExtrapolationOfir Press, Noah A. Smith, Mike LewisICLR 2022 · 1,168 citations
- Mitigating Frequency Bias and Anisotropy in Language Model Pre-Training with Syntactic SmoothingRichard Diehl Martinez, Zébulon Goriely, Andrew Caines, Paula Buttery et al.EMNLP 2024 · 1 citation
- Child-Directed Language Does Not Consistently Boost Syntax Learning in Language ModelsFrancesca Padovani, Jaap Jumelet, Yevgen Matusevych, Arianna BisazzaEMNLP 2025
Related papers
- Working Memory Capacity of ChatGPT: An Empirical StudyDongyu Gong, Xingchen Wan, Dingmin WangAAAI 2024 · 31 citations
- The Rise and Down of Babel Tower: Investigating the Evolution Process of Multilingual Code Large Language ModelJiawei Chen, Wentao Chen, Jing Su, Jingjing Xu et al.ICLR 2025
- SALAAD: Sparse And Low-Rank Adaptation via ADMM for Large Language Model InferenceHao Ma, Melis Ilayda Bal, Liang Zhang, Bingcong Li et al.ICML 2026
- LightThinker: Thinking Step-by-Step CompressionJintian Zhang, Yuqi Zhu, Mengshu Sun, Yujie Luo et al.EMNLP 2025 · 2 citations
- Development of Cognitive Intelligence in Pre-trained Language ModelsRaj Sanjay Shah, Khushi Bhardwaj, Sashank VarmaEMNLP 2024 · 2 citations
