Direct Policy Gradients: Direct Optimization of Policies in Discrete Action Spaces
Guy Lorberbom, Chris J. Maddison, Nicolas Heess, Tamir Hazan, Daniel Tarlow
Abstract
Direct optimization [25, 37] is an appealing framework that replaces integration with optimization of a random objective for approximating gradients in models with discrete random variables [21] . A sampling [24] is a framework for optimizing such random objectives over large spaces. We show how to combine these techniques to yield a reinforcement learning algorithm that approximates a policy gradient by finding trajectories that optimize a random objective. We call the resulting algorithms direct policy gradient (DirPG) algorithms. A main benefit of DirPG algorithms is that they allow the insertion of domain knowledge in the form of upper bounds on return-to-go at training time, like is used in heuristic search, while still directly computing a policy gradient. We further analyze their properties, showing there are cases where DirPG has an exponentially larger probability of sampling informative gradients compared to REINFORCE. We also show that there is a built-in variance reduction technique and that a parameter that was previously viewed as a numerical approximation can be interpreted as controlling risk sensitivity. Empirically, we evaluate the effect of key degrees of freedom and show that the algorithm performs well in illustrative domains compared to baselines. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 1633a0e7-eb7e-4b7a-b196-dd432ccbc5e6Cited by top-tier papers4
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 59 citations
- Automatic Differentiation of Programs with Discrete RandomnessGaurav Arya, Moritz Schauer, Frank Schäfer, Christopher RackauckasNeurIPS 2022 · 56 citations
- Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative TrickLennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos MartiresNeurIPS 2023 · 17 citations
- Learning Randomly Perturbed Structured Predictors for Direct Loss MinimizationHedda Cohen Indelman, Tamir HazanICML 2021 · 9 citations
Builds on1
Related papers
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 67 citations
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 1 citation
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 26 citations
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 4 citations
- Learning to Schedule in Diffusion Probabilistic ModelsYunke Wang, Xiyu Wang, Anh-Dung Dinh, Bo Du et al.KDD 2023 · 17 citations
