Correcting discount-factor mismatch in on-policy policy gradient methods
Fengdi Che, Gautham Vasan, A. Rupam Mahmood
摘要
The policy gradient theorem gives a convenient form of the policy gradient in terms of three factors: an action value, a gradient of the action likelihood, and a state distribution involving discounting called the discounted stationary distribution. But commonly used on-policy methods based on the policy gradient theorem ignores the discount factor in the state distribution, which is technically incorrect and may even cause degenerate learning behavior in some environments. An existing solution corrects this discrepancy by using as a factor in the gradient estimate. However, this solution is not widely adopted and does not work well in tasks where the later states are similar to earlier states. We introduce a novel distribution correction to account for the discounted stationary distribution that can be plugged into many existing gradient estimators. Our correction circumvents the performance degradation associated with the correction with a lower variance. Importantly, compared to the uncorrected estimators, our algorithm provides improved state emphasis to evade suboptimal policies in certain environments and consistently matches or exceeds the original performance on several OpenAI gym and DeepMind suite benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper3
- Deep Policy Gradient Methods Without Batch Updates, Target Networks, or Replay BuffersGautham Vasan, Mohamed Elsayed, Seyed Alireza Azimi, Jiamin He 等NeurIPS 2024 · 被引用 27 次
- Revisiting Scalable Hessian Diagonal Approximations for Applications in Reinforcement LearningMohamed Elsayed, Homayoon Farrahi, Felix Dangel, A. Rupam MahmoodICML 2024 · 被引用 7 次
- Target Networks and Over-parameterization Stabilize Off-policy Bootstrapping with Function ApproximationFengdi Che, Chenjun Xiao, Jincheng Mei, Bo Dai 等ICML 2024 · 被引用 7 次
它引用的顶会 Paper2
相关 Paper
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White 等ICML 2020 · 被引用 72 次
- Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient MethodsXin Guo, Anran Hu, Junzi ZhangAAAI 2022 · 被引用 10 次
- Infinite-horizon Off-Policy Policy Evaluation with Multiple Behavior PoliciesXinyun Chen, Lu Wang, Yizhe Hang, Heng Ge 等ICLR 2020 · 被引用 5 次
- A Temporal-Difference Approach to Policy Gradient EstimationSamuele Tosatto, Andrew Patterson, Martha White, Rupam MahmoodICML 2022 · 被引用 3 次
- Relaxed Stationary Distribution Correction Estimation for Improved Offline Policy OptimizationWoosung Kim, Donghyeon Ki, Byung-Jun LeeAAAI 2024 · 被引用 4 次
