Expected Eligibility Traces
Hado van Hasselt, Sephora Madjiheurem, Matteo Hessel, David Silver, André Barreto, Diana Borsa
摘要
The question of how to determine which states and actions are responsible for a certain outcome is known as the credit assignment problem and remains a central research question in reinforcement learning and artificial intelligence. Eligibility traces enable efficient credit assignment to the recent sequence of states and actions experienced by the agent, but not to counterfactual sequences that could also have led to the current state. In this work, we introduce expected eligibility traces. Expected traces allow, with a single update, to update states and actions that could have preceded the current state, even if they did not do so on this occasion. We discuss when expected traces provide benefits over classic (instantaneous) traces in temporal-difference learning, and show that sometimes substantial improvements can be attained. We provide a way to smoothly interpolate between instantaneous and expected traces by a mechanism similar to bootstrapping, which ensures that the resulting algorithm is a strict generalisation of TD(λ). Finally, we discuss possible extensions and connections to related ideas, such as successor features. Appropriate credit assignment has long been a major research topic in artificial intelligence (Minsky 1963). To make effective decisions and understand the world, we need to accurately associate events, like rewards or penalties, to relevant earlier decisions or situations. This is important both for learning accurate predictions, and for making good decisions. Temporal credit assignment can be achieved with repeated temporal-difference (TD) updates (Sutton 1988). One-step TD updates propagate information slowly: when a surprising value is observed, the state immediately preceding it is updated, but no earlier states or decisions are updated. Multistep updates (Sutton 1988; Sutton and Barto 2018) propagate information faster over longer temporal spans, speeding up credit assignment and learning. Multi-step updates can be implemented online using eligibility traces (Sutton 1988), without incurring significant additional computational expense, even if the time spans are long; these algorithms have computation that is independent of the temporal span of the prediction (van Hasselt and Sutton 2015) . Traces provide temporal credit assignment, but do not assign credit counterfactually to states or actions that could have led to the current state, but did not do so this time.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Forethought and Hindsight in Credit AssignmentVeronica Chelu, Doina Precup, Hado van HasseltNeurIPS 2020 · 被引用 29 次
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay BuffersGautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He 等NeurIPS 2024 · 被引用 27 次
- Maximum State Entropy Exploration using Predecessor and Successor RepresentationsArnav Kumar Jain, Lucas Lehnert, Irina Rish, Glen BersethNeurIPS 2023 · 被引用 27 次
- Emphatic Algorithms for Deep Reinforcement LearningRay Jiang, Tom Zahavy, Zhongwen Xu, Adam White 等ICML 2021 · 被引用 22 次
- Would I have gotten that reward? Long-term credit assignment by counterfactual contribution analysisAlexander Meulemans, Simon Schug, Seijin Kobayashi, Nathaniel D. Daw 等NeurIPS 2023 · 被引用 16 次
它引用的顶会 Paper1
相关 Paper
- From Past to Future: Rethinking Eligibility TracesDhawal Gupta, Scott M. Jordan, Shreyas Chaudhari, Bo Liu 等AAAI 2024 · 被引用 5 次
- A Generalized Bootstrap Target for Value-Learning, Efficiently Combining Value and Feature PredictionsAnthony GX-Chen, Veronica Chelu, Blake A. Richards, Joelle PineauAAAI 2022 · 被引用 1 次
- Sequence Compression Speeds Up Credit Assignment in Reinforcement LearningAditya A. Ramesh, Kenny John Young, Louis Kirsch, Jürgen SchmidhuberICML 2024 · 被引用 2 次
- Learning Guidance Rewards with Trajectory-space SmoothingTanmay Gangwani, Yuan Zhou, Jian PengNeurIPS 2020 · 被引用 46 次
- Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement LearningBrett Daley, Martha White, Christopher Amato, Marlos C. MachadoICML 2023 · 被引用 4 次
