A Local Temporal Difference Code for Distributional Reinforcement Learning
Pablo Tano, Peter Dayan, Alexandre Pouget
Abstract
Recent theoretical and experimental results suggest that the dopamine system implements distributional temporal difference backups, allowing learning of the entire distributions of the long-run values of states rather than just their expected values. However, the distributional codes explored so far rely on a complex imputation step which crucially relies on spatial non-locality: in order to compute reward prediction errors, units must know not only their own state but also the states of the other units. It is far from clear how these steps could be implemented in realistic neural circuits. Here, we propose a local temporal difference code for distributional reinforcement learning that is representationally powerful and computationally straightforward. The code decomposes value distributions and prediction errors across three completely separated dimensions: reward magnitude (related to distributional quantiles), time horizon (related to eligibility traces) and temporal discounting (related to the Laplace transform of future immediate rewards). Besides lending itself to a local learning rule, the decomposition can be exploited by model-based computations, for instance allowing immediate adjustments to changing horizons or discount factors. Finally, we show that our code can be computed linearly from an ensemble of successor representations with multiple temporal discounts which, according to a recent proposal, might be implemented in the hippocampus.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d08ee0dd-d02f-408f-9ef0-21debc57634cCited by top-tier papers3
- Beyond accuracy: generalization properties of bio-plausible temporal credit assignment rulesYuhan Helena Liu, Arna Ghosh, Blake A. Richards, Eric Shea-Brown et al.NeurIPS 2022 · 10 citations
- Distributional Bellman Operators over Mean EmbeddingsLi Kevin Wenliang, Grégoire Delétang, Matthew Aitchison, Marcus Hutter et al.ICML 2024 · 5 citations
- Multi-timescale Reinforcement Learning by Value ReconstructionZhan Su, Peixi Peng, Xinyu Hu, Cong Li et al.ICML 2026
Related papers
- Temporal-Difference Learning Using Distributed Error SignalsJonas Guan, Shon Eduard Verch, Claas Voelcker, Ethan C. Jackson et al.NeurIPS 2024 · 5 citations
- A Differential Perspective on Distributional Reinforcement LearningJuan Sebastian Rojas, Chi-Guhn LeeAAAI 2026 · 4 citations
- Distributional Reinforcement Learning for Multi-Dimensional Reward FunctionsPushi Zhang, Xiaoyu Chen, Li Zhao, Wei Xiong et al.NeurIPS 2021 · 33 citations
- Predictive auxiliary objectives in deep RL mimic learning in the brainChing Fang, Kim StachenfeldICLR 2024 · 16 citations
- The Statistical Benefits of Quantile Temporal-Difference Learning for Value EstimationMark Rowland, Yunhao Tang, Clare Lyle, Rémi Munos et al.ICML 2023 · 13 citations
