Understanding Policy Gradient Algorithms: A Sensitivity-Based Approach
Shuang Wu, Ling Shi, Jun Wang, Guangjian Tian
摘要
The REINFORCE algorithm from Williams is popular in policy gradient (PG) for solving reinforcement learning (RL) problems. Meanwhile, the theoretical form of PG is from Sutton et al. Although both formulae prescribe PG, their precise connections are not yet illustrated. Recently, Nota and Thomas (2020) have found that the ambiguity causes implementation errors. Motivated by the ambiguity and implementation incorrectness, we study PG from a perturbation perspective. In particular, we derive PG in a unified framework, precisely clarify the relation between PG implementation and theory, and echo back the findings by Nota and Thomas. Diving into factors contributing to empirical successes of the existing erroneous implementations, we find that small approximation error and the experience replay mechanism play critical roles. 1 This issue is so widely spread that even influential platforms, such as OpenAI Spinning Up https:// spinningup.openai.com/en/latest/index.html and MATLAB toolbox https://www.mathworks.com/ help/reinforcement-learning/agents.html , inherit this issue.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- On the Second-Order Convergence of Biased Policy Gradient AlgorithmsSiqiao Mu, Diego KlabjanICML 2024 · 被引用 4 次
- On Many-Actions Policy GradientMichal Nauman, Marek CyganICML 2023 · 被引用 1 次
它引用的顶会 Paper7
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Implementation Matters in Deep RL: A Case Study on PPO and TRPOLogan Engstrom, Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 305 次
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 被引用 162 次
- A Closer Look at Deep Policy GradientsAndrew Ilyas, Logan Engstrom, Shibani Santurkar, Dimitris Tsipras 等ICLR 2020 · 被引用 107 次
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári 等NeurIPS 2021 · 被引用 87 次
相关 Paper
- On the Mistaken Assumption of Interchangeable Deep Reinforcement Learning ImplementationsRajdeep Singh Hundal, Yan Xiao, Xiaochun Cao, Jin Song Dong 等ICSE 2025
- Performative Policy Gradient: Optimality in Performative Reinforcement LearningDebabrota Basu, Udvas Das, Brahim Driss, Uddalak MukherjeeICML 2026 · 被引用 2 次
- Characterizing the Gap Between Actor-Critic and Policy GradientJunfeng Wen, Saurabh Kumar, Ramki Gummadi, Dale SchuurmansICML 2021 · 被引用 18 次
- An operator view of policy gradient methodsDibya Ghosh, Marlos C. Machado, Nicolas Le RouxNeurIPS 2020 · 被引用 30 次
- Settling the Variance of Multi-Agent Policy GradientsJakub Grudzien Kuba, Muning Wen, Linghui Meng, Shangding Gu 等NeurIPS 2021 · 被引用 121 次
