Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement Learning
Brett Daley, Martha White, Christopher Amato, Marlos C. Machado
Abstract
Off-policy learning from multistep returns is crucial for sample-efficient reinforcement learning, but counteracting off-policy bias without exacerbating variance is challenging. Classically, offpolicy bias is corrected in a per-decision manner: past temporal-difference errors are re-weighted by the instantaneous Importance Sampling (IS) ratio after each action via eligibility traces. Many offpolicy algorithms rely on this mechanism, along with differing protocols for cutting the IS ratios to combat the variance of the IS estimator. Unfortunately, once a trace has been fully cut, the effect cannot be reversed. This has led to the development of credit-assignment strategies that account for multiple past experiences at a time. These trajectory-aware methods have not been extensively analyzed, and their theoretical justification remains uncertain. In this paper, we propose a multistep operator that can express both perdecision and trajectory-aware methods. We prove convergence conditions for our operator in the tabular setting, establishing the first guarantees for several existing methods as well as many new ones. Finally, we introduce Recency-Bounded Importance Sampling (RBIS), which leverages trajectory awareness to perform robustly across λ-values in several off-policy control tasks. Introduction Reinforcement learning concerns an agent interacting with its environment through trial and error to maximize its expected cumulative reward. One of the great challenges of re-
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 06e0a9c3-e5ee-42ba-8cae-6fe542bdb7c9Cited by top-tier papers3
- Q-learning Decision Transformer: Leveraging Dynamic Programming for Conditional Sequence Modelling in Offline RLTaku Yamagata, Ahmed Khalil, Raúl Santos-RodríguezICML 2023 · 121 citations
- Averaging n-step Returns Reduces Variance in Reinforcement LearningBrett Daley, Martha White, Marlos C. MachadoICML 2024 · 7 citations
- Peng's Q(π) for Conservative Value Estimation in Offline Reinforcement LearningByeongchan Kim, Min-hwan OhICLR 2026
Builds on1
Related papers
- Finite-Sample Analysis of Off-Policy TD-Learning via Generalized Bellman OperatorsZaiwei Chen, Siva Theja Maguluri, Sanjay Shakkottai, Karthikeyan ShanmugamNeurIPS 2021 · 21 citations
- Reusing Trajectories in Policy Gradients Enables Fast ConvergenceAlessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini et al.ICML 2026
- Learning to Sample with Local and Global Contexts in Experience Replay BufferYoungmin Oh, Kimin Lee, Jinwoo Shin, Eunho Yang et al.ICLR 2021 · 19 citations
- Demystifying the Paradox of Importance Sampling with an Estimated History-Dependent Behavior Policy in Off-Policy EvaluationHongyi Zhou, Josiah P. Hanna, Jin Zhu, Ying Yang et al.ICML 2025
- Backstepping Temporal Difference LearningHan-Dong Lim, Donghwan LeeICLR 2023
