Temporal-Difference Learning Using Distributed Error Signals
Jonas Guan, Shon Eduard Verch, Claas Voelcker, Ethan C. Jackson, Nicolas Papernot, William A. Cunningham
摘要
A computational problem in biological reward-based learning is how credit assignment is performed in the nucleus accumbens (NAc). Much research suggests that NAc dopamine encodes temporal-difference (TD) errors for learning value predictions. However, dopamine is synchronously distributed in regionally homogeneous concentrations, which does not support explicit credit assignment (like used by backpropagation). It is unclear whether distributed errors alone are sufficient for synapses to make coordinated updates to learn complex, nonlinear reward-based learning tasks. We design a new deep Q-learning algorithm, Artificial Dopamine, to computationally demonstrate that synchronously distributed, per-layer TD errors may be sufficient to learn surprisingly complex RL tasks. We empirically evaluate our algorithm on MinAtar, the DeepMind Control Suite, and classic control tasks, and show it often achieves comparable performance to deep RL algorithms that use backpropagation.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Identifiable Token Correspondence for World ModelsYoungin Kim, Ray Sun, Inho Kim, Bumsoo Park 等ICML 2026
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-FunctionsZequan Wu, Mengye RenICLR 2026
它引用的顶会 Paper8
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville 等NeurIPS 2021 · 被引用 1,067 次
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 被引用 388 次
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 被引用 125 次
- Error-driven Input Modulation: Solving the Credit Assignment Problem without a Backward PassGiorgia Dellaferrera, Gabriel KreimanICML 2022 · 被引用 80 次
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli PoliciesTim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato 等NeurIPS 2021 · 被引用 59 次
相关 Paper
- A Local Temporal Difference Code for Distributional Reinforcement LearningPablo Tano, Peter Dayan, Alexandre PougetNeurIPS 2020 · 被引用 27 次
- Minimizing Control for Credit Assignment with Strong FeedbackAlexander Meulemans, Matilde Tristany Farinha, Maria R. Cervera, João Sacramento 等ICML 2022 · 被引用 24 次
- Single-phase deep learning in cortico-cortical networksWill Greedy, Heng Wei Zhu, Joseph Pemberton, Jack Mellor 等NeurIPS 2022 · 被引用 62 次
- Biologically-plausible backpropagation through arbitrary timespans via local neuromodulatorsYuhan Helena Liu, Stephen Smith, Stefan Mihalas, Eric Shea-Brown 等NeurIPS 2022 · 被引用 18 次
- Biological credit assignment through dynamic inversion of feedforward networksWilliam F. Podlaski, Christian K. MachensNeurIPS 2020 · 被引用 26 次
