Predictive auxiliary objectives in deep RL mimic learning in the brain
Ching Fang, Kim Stachenfeld
摘要
The ability to predict upcoming events has been hypothesized to comprise a key aspect of natural and machine cognition. This is supported by trends in deep reinforcement learning (RL), where self-supervised auxiliary objectives such as prediction are widely used to support representation learning and improve task performance. Here, we study the effects predictive auxiliary objectives have on representation learning across different modules of an RL system and how these mimic representational changes observed in the brain. We find that predictive objectives improve and stabilize learning particularly in resource-limited architectures, and we identify settings where longer predictive horizons better support representational transfer. Furthermore, we find that representational changes in this RL system bear a striking resemblance to changes in neural activity observed in the brain across various experiments. Specifically, we draw a connection between the auxiliary predictive model of the RL system and hippocampus, an area thought to learn a predictive model to support memory-guided behavior. We also connect the encoder network and the value learning network of the RL system to visual cortex and striatum in the brain, respectively. This work demonstrates how representation learning in deep RL systems can provide an interpretable framework for modeling multi-region interactions in the brain. The deep RL perspective taken here also suggests an additional role of the hippocampus in the brain -- that of an auxiliary learning system that benefits representation learning in other regions.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Brain-Like Processing Pathways Form in Models With Heterogeneous ExpertsJack Cook, Danyal Akarca, Rui Ponte Costa, Jascha AchterbergNeurIPS 2025 · 被引用 6 次
- A Model of Place Field Reorganization During Reward MaximizationM. Ganesh Kumar, Blake Bordelon, Jacob A. Zavatone-Veth, Cengiz PehlevanICML 2025
- Predictive Coding Enhances Meta-RL To Achieve Interpretable Bayes-Optimal Belief Representation Under Partial ObservabilityPo-Chen Kuo, Han Hou, Will Dabney, Edgar Y. WalkerNeurIPS 2025
- From Observations to Events: Event-Aware World Models for Reinforcement LearningZhao-Han Peng, Shaohui Li, Zhi Li, Shulan Ruan 等ICLR 2026
它引用的顶会 Paper10
- Fast Task Inference with Variational Intrinsic Successor FeaturesSteven Hansen, Will Dabney, André Barreto, David Warde-Farley 等ICLR 2020 · 被引用 176 次
- On the role of planning in model-based deep reinforcement learningJessica B. Hamrick, Abram L. Friesen, Feryal M. P. Behbahani, Arthur Guez 等ICLR 2021 · 被引用 77 次
- The Value-Improvement Path: Towards Better Representations for Reinforcement LearningWill Dabney, André Barreto, Mark Rowland, Robert Dadashi 等AAAI 2021 · 被引用 76 次
- Novelty Search in Representational Space for Sample Efficient ExplorationRuo Yu Tao, Vincent François-Lavet, Joelle PineauNeurIPS 2020 · 被引用 53 次
- Understanding Self-Predictive Learning for Reinforcement LearningYunhao Tang, Zhaohan Daniel Guo, Pierre Harvey Richemond, Bernardo Ávila Pires 等ICML 2023 · 被引用 46 次
相关 Paper
- A Multi-Region Brain Model to Elucidate the Role of Hippocampus in Spatially Embedded Decision-MakingYi Xie, Jaedong Hwang, Carlos D. Brody, David W. Tank 等ICML 2025
- The functional specialization of visual cortex emerges from training parallel pathways with self-supervised predictive learningShahab Bakhtiari, Patrick J. Mineault, Timothy P. Lillicrap, Christopher C. Pack 等NeurIPS 2021 · 被引用 103 次
- A Cognitive Model for Learning Abstract Relational Structures from Memory-based Decision-Making TasksHaruo HosoyaICLR 2024 · 被引用 1 次
- On the Unexpected Effectiveness of Reinforcement Learning for Sequential RecommendationAlvaro Labarca, Denis Parra, Rodrigo Toro IcarteICML 2024 · 被引用 1 次
- Reinforcement Learning based Disease Progression Model for Alzheimer's DiseaseKrishnakant V. Saboo, Anirudh Choudhary, Yurui Cao, Gregory A. Worrell 等NeurIPS 2021 · 被引用 19 次
