Lune

ICML2022Top-tier venue

Understanding Policy Gradient Algorithms: A Sensitivity-Based Approach

Shuang Wu, Ling Shi, Jun Wang, Guangjian Tian

2022Year
7Citations
2Top-tier citations

Abstract

The REINFORCE algorithm from Williams is popular in policy gradient (PG) for solving reinforcement learning (RL) problems. Meanwhile, the theoretical form of PG is from Sutton et al. Although both formulae prescribe PG, their precise connections are not yet illustrated. Recently, Nota and Thomas (2020) have found that the ambiguity causes implementation errors. Motivated by the ambiguity and implementation incorrectness, we study PG from a perturbation perspective. In particular, we derive PG in a unified framework, precisely clarify the relation between PG implementation and theory, and echo back the findings by Nota and Thomas. Diving into factors contributing to empirical successes of the existing erroneous implementations, we find that small approximation error and the experience replay mechanism play critical roles. 1 This issue is so widely spread that even influential platforms, such as OpenAI Spinning Up https:// spinningup.openai.com/en/latest/index.html and MATLAB toolbox https://www.mathworks.com/ help/reinforcement-learning/agents.html , inherit this issue.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext fc169aa7-e1cf-4a14-8ee3-b0576ad71b6a

Cited by top-tier papers2

Ask how each one uses it

Builds on7

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines