Modelling the control of offline processing with reinforcement learning
Eleanor Spens, Neil Burgess, Tim E. J. Behrens
Abstract
Brains reorganise knowledge offline to improve future behaviour, with 'replay' involved in consolidating memories, abstracting patterns from experience, and simulating new scenarios. However, there are few models of how the brain might orchestrate these processes, and of when different types of replay might be useful.
Here we propose a framework in which a meta-controller learns to coordinate offline learning of a lower-level agent or model in 'sleep' phases to maximise reward in a 'wake' phase. The meta-controller selects among several actions, such as learning from recent memories in a hippocampal store, abstracting patterns from memories into a 'world model', and learning from generated data. In addition, the meta-controller learns to estimate the value of each episode, enabling the prioritisation of past events in memory replay, or of new simulations in generative replay. Using image classification, maze solving, and relational inference tasks, we show that the meta-controller learns an adaptive curriculum for offline learning. This lays the groundwork for normative predictions about replay in a range of experimental neuroscience tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on2
Related papers
- Learning Human Habits with Rule-Guided Active InferenceGong Zhiren, Chao Yang, Wendi Ren, Shuang LiICLR 2026
- Biologically inspired sleep algorithm for increased generalization and adversarial robustness in deep neural networksTimothy Tadros, Giri P. Krishnan, Ramyaa Ramyaa, Maxim BazhenovICLR 2020 · 28 citations
- Robust Task Representations for Offline Meta-Reinforcement Learning via Contrastive LearningHaoqi Yuan, Zongqing LuICML 2022 · 53 citations
- Fast And Slow Learning Of Recurrent Independent MechanismsKanika Madan, Nan Rosemary Ke, Anirudh Goyal, Bernhard Schölkopf et al.ICLR 2021 · 41 citations
- Policy Rehearsing: Training Generalizable Policies for Reinforcement LearningChengxing Jia, Chenxiao Gao, Hao Yin, Fuxiang Zhang et al.ICLR 2024 · 6 citations
