Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
Brett Daley, Martha White, Christopher Amato, Marlos C. Machado
摘要
Off-policy learning from multistep returns is crucial for sample-efficient reinforcement learning, but counteracting off-policy bias without exacerbating variance is challenging. Classically, offpolicy bias is corrected in a per-decision manner: past temporal-difference errors are re-weighted by the instantaneous Importance Sampling (IS) ratio after each action via eligibility traces. Many offpolicy algorithms rely on this mechanism, along with differing protocols for cutting the IS ratios to combat the variance of the IS estimator. Unfortunately, once a trace has been fully cut, the effect cannot be reversed. This has led to the development of credit-assignment strategies that account for multiple past experiences at a time. These trajectory-aware methods have not been extensively analyzed, and their theoretical justification remains uncertain. In this paper, we propose a multistep operator that can express both perdecision and trajectory-aware methods. We prove convergence conditions for our operator in the tabular setting, establishing the first guarantees for several existing methods as well as many new ones. Finally, we introduce Recency-Bounded Importance Sampling (RBIS), which leverages trajectory awareness to perform robustly across λ-values in several off-policy control tasks. Introduction Reinforcement learning concerns an agent interacting with its environment through trial and error to maximize its expected cumulative reward. One of the great challenges of re-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 被引用 121 次
- Averaging n-step Returns Reduces Variance in Reinforcement LearningBrett Daley, Martha White, Marlos C. MachadoICML 2024 · 被引用 7 次
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
它引用的顶会 Paper1
相关 Paper
- Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman OperatorsZaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, Karthikeyan ShanmugamNeurIPS 2021 · 被引用 21 次
- Reusing Trajectories in Policy Gradients Enables Fast ConvergenceAlessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini 等ICML 2026
- Learning to Sample with Local and Global Contexts in Experience Replay BufferYoungmin Oh, Kimin Lee, Jinwoo Shin, Eunho Yang 等ICLR 2021 · 被引用 19 次
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang 等ICML 2025
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
