Modelling the control of offline processing with reinforcement learning
Eleanor Spens, Neil Burgess, Tim E. J. Behrens
摘要
Brains reorganise knowledge offline to improve future behaviour, with 'replay' involved in consolidating memories, abstracting patterns from experience, and simulating new scenarios. However, there are few models of how the brain might orchestrate these processes, and of when different types of replay might be useful.
Here we propose a framework in which a meta-controller learns to coordinate offline learning of a lower-level agent or model in 'sleep' phases to maximise reward in a 'wake' phase. The meta-controller selects among several actions, such as learning from recent memories in a hippocampal store, abstracting patterns from memories into a 'world model', and learning from generated data. In addition, the meta-controller learns to estimate the value of each episode, enabling the prioritisation of past events in memory replay, or of new simulations in generative replay. Using image classification, maze solving, and relational inference tasks, we show that the meta-controller learns an adaptive curriculum for offline learning. This lays the groundwork for normative predictions about replay in a range of experimental neuroscience tasks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Learning Human Habits with Rule-Guided Active InferenceGong Zhiren, Chao Yang, Wendi Ren, Shuang LiICLR 2026
- Biologically inspired sleep algorithm for increased generalization and adversarial robustness in deep neural networksTimothy Tadros, Giri P. Krishnan, Ramyaa Ramyaa, Maxim BazhenovICLR 2020 · 被引用 28 次
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 被引用 53 次
- Fast And Slow Learning Of Recurrent Independent MechanismsKanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf 等ICLR 2021 · 被引用 41 次
- Policy Rehearsing: Training Generalizable Policies for Reinforcement LearningChengxing Jia, Chenxiao Gao, Hao Yin, Fuxiang Zhang 等ICLR 2024 · 被引用 6 次
