Structural Credit Assignment in Neural Networks using Reinforcement Learning
Dhawal Gupta, Gabor Mihucz, Matthew Schlegel, James E. Kostas, Philip S. Thomas, Martha White
Abstract
Structural credit assignment in neural networks is a long-standing problem, with a variety of alternatives to backpropagation proposed to allow for local training of nodes. One of the early strategies was to treat each node as an agent and use a reinforcement learning method called REINFORCE to update each node locally with only a global reward signal. In this work, we revisit this approach and investigate if we can leverage other reinforcement learning approaches to improve learning. We first formalize training a neural network as a finite-horizon reinforcement learning problem and discuss how this facilitates using ideas from reinforcement learning like off-policy learning, exploration and planning. We first show that the standard REINFORCE approach can learn but is suboptimal due to on-policy training: each agent learns to output an activation under suboptimal action selection from the other agents. We show that we can overcome this suboptimality with an off-policy approach, that it is particularly effective with discretized actions. We provide several additional experiments, highlighting the utility of exploration, robustness to correlated samples when learning online and a study into the policy parameterization of each agent.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 240f0bbf-32ce-42e0-8192-922239dccfe9Cited by top-tier papers2
- When Can Model-Free Reinforcement Learning be Enough for Thinking?Josiah Hanna, Nicholas CorradoNeurIPS 2025 · 6 citations
- Hexaïssa: Standing on Giants' Shoulders - Routing the Best Chess Engines with Mixture-of-Experts and Latent Reward LearningBach Ngo, Nguyen Hoang Khoi DoAAAI 2026
Builds on7
- PC-PG: Policy Cover Directed Exploration for Provable Policy Gradient LearningAlekh Agarwal, Mikael Henaff, Sham M. Kakade, Wen SunNeurIPS 2020 · 126 citations
- Leverage the Average: an Analysis of KL Regularization in Reinforcement LearningNino Vieillard, Tadashi Kozuno, Bruno Scherrer, Olivier Pietquin et al.NeurIPS 2020 · 106 citations
- Learning to solve the credit assignment problemBenjamin James Lansdell, Prashanth Ravi Prakash, Konrad Paul KördingICLR 2020 · 60 citations
- Fuzzy Tiling Activations: A Simple Approach to Learning Sparse Representations OnlineYangchen Pan, Kirby Banman, Martha WhiteICLR 2021 · 24 citations
- Path Sample-Analytic Gradient Estimators for Stochastic Binary NetworksAlexander Shekhovtsov, Viktor Yanush, Boris FlachNeurIPS 2020 · 14 citations
Related papers
- Learning by Competition of Self-Interested Reinforcement Learning AgentsStephen ChungAAAI 2022 · 5 citations
- Asynchronous Coagent NetworksJames E. Kostas, Chris Nota, Philip S. ThomasICML 2020 · 9 citations
- MAP Propagation Algorithm: Faster Learning with a Team of Reinforcement Learning AgentsStephen ChungNeurIPS 2021 · 5 citations
- Hindsight Network Credit Assignment: Efficient Credit Assignment in Networks of Discrete Stochastic UnitsKenny YoungAAAI 2022
- Learning to Learn with Feedback and Local PlasticityJack Lindsey, Ashok Litwin-KumarNeurIPS 2020 · 38 citations
