Temporal-Difference Learning Using Distributed Error Signals
Jonas Guan, Shon Eduard Verch, Claas Voelcker, Ethan C. Jackson, Nicolas Papernot, William A. Cunningham
Abstract
A computational problem in biological reward-based learning is how credit assignment is performed in the nucleus accumbens (NAc). Much research suggests that NAc dopamine encodes temporal-difference (TD) errors for learning value predictions. However, dopamine is synchronously distributed in regionally homogeneous concentrations, which does not support explicit credit assignment (like used by backpropagation). It is unclear whether distributed errors alone are sufficient for synapses to make coordinated updates to learn complex, nonlinear reward-based learning tasks. We design a new deep Q-learning algorithm, Artificial Dopamine, to computationally demonstrate that synchronously distributed, per-layer TD errors may be sufficient to learn surprisingly complex RL tasks. We empirically evaluate our algorithm on MinAtar, the DeepMind Control Suite, and classic control tasks, and show it often achieves comparable performance to deep RL algorithms that use backpropagation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Identifiable Token Correspondence for World ModelsYoungin Kim, Ray Sun, Inho Kim, Bumsoo Park et al.ICML 2026
- Local Reinforcement Learning with Action-Conditioned Root Mean Squared Q-FunctionsZequan Wu, Mengye RenICLR 2026
Builds on8
- Deep Reinforcement Learning at the Edge of the Statistical PrecipiceRishabh Agarwal, Max Schwarzer, Pablo Samuel Castro, Aaron C. Courville et al.NeurIPS 2021 · 1,067 citations
- TD-MPC2: Scalable, Robust World Models for Continuous ControlNicklas Hansen, Hao Su, Xiaolong WangICLR 2024 · 388 citations
- Revisiting Rainbow: Promoting more insightful and inclusive deep reinforcement learning researchJohan S. Obando-Ceron, Pablo Samuel CastroICML 2021 · 125 citations
- Error-driven Input Modulation: Solving the Credit Assignment Problem without a Backward PassGiorgia Dellaferrera, Gabriel KreimanICML 2022 · 80 citations
- Is Bang-Bang Control All You Need? Solving Continuous Control with Bernoulli PoliciesTim Seyde, Igor Gilitschenski, Wilko Schwarting, Bartolomeo Stellato et al.NeurIPS 2021 · 59 citations
Related papers
- A Local Temporal Difference Code for Distributional Reinforcement LearningPablo Tano, Peter Dayan, Alexandre PougetNeurIPS 2020 · 27 citations
- Minimizing Control for Credit Assignment with Strong FeedbackAlexander Meulemans, Matilde Tristany Farinha, Maria R. Cervera, João Sacramento et al.ICML 2022 · 24 citations
- Single-phase deep learning in cortico-cortical networksWill Greedy, Heng Wei Zhu, Joseph Pemberton, Jack Mellor et al.NeurIPS 2022 · 62 citations
- Biologically-plausible backpropagation through arbitrary timespans via local neuromodulatorsYuhan Helena Liu, Stephen Smith, Stefan Mihalas, Eric Shea-Brown et al.NeurIPS 2022 · 18 citations
- Biological credit assignment through dynamic inversion of feedforward networksWilliam F. Podlaski, Christian K. MachensNeurIPS 2020 · 26 citations
