Correcting discount-factor mismatch in on-policy policy gradient methods
Fengdi Che, Gautham Vasan, A. Rupam Mahmood
Abstract
The policy gradient theorem gives a convenient form of the policy gradient in terms of three factors: an action value, a gradient of the action likelihood, and a state distribution involving discounting called the discounted stationary distribution. But commonly used on-policy methods based on the policy gradient theorem ignores the discount factor in the state distribution, which is technically incorrect and may even cause degenerate learning behavior in some environments. An existing solution corrects this discrepancy by using as a factor in the gradient estimate. However, this solution is not widely adopted and does not work well in tasks where the later states are similar to earlier states. We introduce a novel distribution correction to account for the discounted stationary distribution that can be plugged into many existing gradient estimators. Our correction circumvents the performance degradation associated with the correction with a lower variance. Importantly, compared to the uncorrected estimators, our algorithm provides improved state emphasis to evade suboptimal policies in certain environments and consistently matches or exceeds the original performance on several OpenAI gym and DeepMind suite benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext ca04c695-10cb-43f9-b57b-d02fad0edb5aCited by top-tier papers3
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay BuffersGautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He et al.NeurIPS 2024 · 27 citations
- Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement LearningMohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam MahmoodICML 2024 · 7 citations
- Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function ApproximationFengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai et al.ICML 2024 · 7 citations
Builds on2
Related papers
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient MethodsXin Guo, Anran Hu, Junzi ZhangAAAI 2022 · 10 citations
- Infinite-horizon Off-Policy Policy Evaluation with Multiple Behavior PoliciesXinyun Chen, Lu Wang, Yizhe Hang, Heng Ge et al.ICLR 2020 · 5 citations
- A Temporal-Difference Approach to Policy Gradient EstimationSamuele Tosatto, Andrew Patterson, Martha White, Rupam MahmoodICML 2022 · 3 citations
- Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy OptimizationWoosung Kim, Donghyeon Ki, Byung-Jun LeeAAAI 2024 · 4 citations
