ICML2022

Understanding Policy Gradient Algorithms: A Sensitivity-Based Approach

Shuang Wu, Ling Shi, Jun Wang, Guangjian Tian

7 citations

Abstract

The REINFORCE algorithm from Williams is popular in policy gradient (PG) for solving reinforcement learning (RL) problems. Meanwhile, the theoretical form of PG is from Sutton et al. Although both formulae prescribe PG, their precise connections are not yet illustrated. Recently, Nota and Thomas (2020) have found that the ambiguity causes implementation errors. Motivated by the ambiguity and implementation incorrectness, we study PG from a perturbation perspective. In particular, we derive PG in a unified framework, precisely clarify the relation between PG implementation and theory, and echo back the findings by Nota and Thomas. Diving into factors contributing to empirical successes of the existing erroneous implementations, we find that small approximation error and the experience replay mechanism play critical roles. 1 This issue is so widely spread that even influential platforms, such as OpenAI Spinning Up https:// spinningup.openai.com/en/latest/index.html and MATLAB toolbox https://www.mathworks.com/ help/reinforcement-learning/agents.html , inherit this issue.