Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysis
Alexander Meulemans, Simon Schug, Seijin Kobayashi, Nathaniel D. Daw, Gregory Wayne
Abstract
To make reinforcement learning more sample efficient, we need better credit assignment methods that measure an action's influence on future rewards. Building upon Hindsight Credit Assignment (HCA) [1], we introduce Counterfactual Contribution Analysis (COCOA), a new family of model-based credit assignment algorithms. Our algorithms achieve precise credit assignment by measuring the contribution of actions upon obtaining subsequent rewards, by quantifying a counterfactual query: 'Would the agent still have reached this reward if it had taken another action?'. We show that measuring contributions w.r.t. rewarding states, as is done in HCA, results in spurious estimates of contributions, causing HCA to degrade towards the high-variance REINFORCE estimator in many relevant environments. Instead, we measure contributions w.r.t. rewards or learned representations of the rewarding objects, resulting in gradient estimates with lower variance. We run experiments on a suite of problems specifically designed to evaluate long-term credit assignment capabilities. By using dynamic programming, we measure ground-truth policy gradients and show that the improved performance of our new model-based credit assignment methods is due to lower bias and variance compared to HCA and common baselines. Our results demonstrate how modeling action contributions towards rewarding outcomes can be leveraged for credit assignment, opening a new path towards sample-efficient reinforcement learning. 2 * Equal contribution; ordering determined by coin flip. 2 Code available at https://github.com/seijin-kobayashi/cocoa 37th Conference on Neural Information Processing Systems (NeurIPS 2023).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 66a3dae1-e510-4379-ad2d-6e557c40fb7dCited by top-tier papers3
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng et al.NeurIPS 2025 · 7 citations
- Sequence Compression Speeds Up Credit Assignment in Reinforcement LearningAditya A. Ramesh, Kenny John Young, Louis Kirsch, Jürgen SchmidhuberICML 2024 · 2 citations
- Predict and Resist: Long-Term Accident Anticipation Under Sensor NoiseXingcheng Liu, Bin Rao, Yanchen Guan, Chengyue Wang et al.AAAI 2026
Builds on14
- Decision Transformer: Reinforcement Learning via Sequence ModelingLili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee et al.NeurIPS 2021 · 2,557 citations
- Dream to Control: Learning Behaviors by Latent ImaginationDanijar Hafner, Timothy P. Lillicrap, Jimmy Ba, Mohammad NorouziICLR 2020 · 1,852 citations
- Mastering Atari with Discrete World ModelsDanijar Hafner, Timothy P. Lillicrap, Mohammad Norouzi, Jimmy BaICLR 2021 · 1,170 citations
- Offline Reinforcement Learning as One Big Sequence Modeling ProblemMichael Janner, Qiyang Li, Sergey LevineNeurIPS 2021 · 950 citations
- Estimating counterfactual treatment outcomes over time through adversarially balanced representationsIoana Bica, Ahmed M. Alaa, James Jordon, Mihaela van der SchaarICLR 2020 · 224 citations
Related papers
- Counterfactual Credit Assignment in Model-Free Reinforcement LearningThomas Mesnard, Theophane Weber, Fabio Viola, Shantanu Thakoor et al.ICML 2021 · 70 citations
- Quantile Credit AssignmentThomas Mesnard, Wenqi Chen, Alaa Saade, Yunhao Tang et al.ICML 2023 · 3 citations
- Hindsight PRIORs for Reward Learning from Human PreferencesMudit Verma, Katherine MetcalfICLR 2024 · 11 citations
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 46 citations
- MARS-RA: Rank Aggregation for Credit Assignment via Multimodal Comparisons in Embodied Multi-Agent CooperationDawei Wang, Di Zhao, Xinyuan Liu, Marci Chi Ma et al.ACL 2026
