Reusing Trajectories in Policy Gradients Enables Fast Convergence
Alessandro Montenegro, Federico Mansutti, Marco Mussi, Matteo Papini, Alberto Maria Metelli
Abstract
Policy gradient (PG) methods are a class of effective reinforcement learning algorithms, particularly when dealing with continuous control problems. They rely on fresh on-policy data, making them sample-inefficient and requiring trajectories to reach an -approximate stationary point. A common strategy to improve efficiency is to reuse information from past iterations, such as previous gradients or trajectories, leading to off-policy PG methods. While gradient reuse has received substantial attention, leading to improved rates up to , the reuse of past trajectories, although intuitive, remains largely unexplored from a theoretical perspective. In this work, we provide the first rigorous theoretical evidence that reusing past off-policy trajectories can significantly accelerate PG convergence. We propose RT-PG (Reusing Trajectories - Policy Gradient), a novel algorithm that leverages a power mean-corrected multiple importance weighting estimator to effectively combine on-policy and off-policy data coming from the most recent iterations. Through a novel analysis, we prove that RT-PG achieves a sample complexity of . When reusing all available past trajectories, this leads to a rate of , the best known one in the literature for PG methods. We further validate our approach empirically, demonstrating its effectiveness against baselines with state-of-the-art rates.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Builds on4
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 99 citations
- Subgaussian and Differentiable Importance Sampling for Off-Policy Evaluation and LearningAlberto Maria Metelli, Alessio Russo, Marcello RestelliNeurIPS 2021 · 55 citations
- PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient EstimationMatilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler H. Summers et al.ICML 2022 · 17 citations
- Learning Optimal Deterministic Policies with Stochastic Policy GradientsAlessandro Montenegro, Marco Mussi, Alberto Maria Metelli, Matteo PapiniICML 2024 · 11 citations
Related papers
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári et al.NeurIPS 2021 · 87 citations
- Trajectory-Aware Eligibility Traces for Off-Policy Reinforcement LearningBrett Daley, Martha White, Christopher Amato, Marlos C. MachadoICML 2023 · 4 citations
- Generalized Proximal Policy Optimization with Sample ReuseJames Queeney, Yannis Paschalidis, Christos G. CassandrasNeurIPS 2021 · 80 citations
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 26 citations
- Efficient Online Reinforcement Learning with Offline DataPhilip J. Ball, Laura Smith, Ilya Kostrikov, Sergey LevineICML 2023 · 326 citations
