Direct Policy Gradients: Direct Optimization of Policies in Discrete Action Spaces
Guy Lorberbom, Chris J. Maddison, Nicolas Heess, Tamir Hazan, Daniel Tarlow
摘要
Direct optimization [25, 37] is an appealing framework that replaces integration with optimization of a random objective for approximating gradients in models with discrete random variables [21] . A sampling [24] is a framework for optimizing such random objectives over large spaces. We show how to combine these techniques to yield a reinforcement learning algorithm that approximates a policy gradient by finding trajectories that optimize a random objective. We call the resulting algorithms direct policy gradient (DirPG) algorithms. A main benefit of DirPG algorithms is that they allow the insertion of domain knowledge in the form of upper bounds on return-to-go at training time, like is used in heuristic search, while still directly computing a policy gradient. We further analyze their properties, showing there are cases where DirPG has an exponentially larger probability of sampling informative gradients compared to REINFORCE. We also show that there is a built-in variance reduction technique and that a parameter that was previously viewed as a numerical approximation can be interpreted as controlling risk sensitivity. Empirically, we evaluate the effect of key degrees of freedom and show that the algorithm performs well in illustrative domains compared to baselines. 34th Conference on Neural Information Processing Systems (NeurIPS 2020),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Estimating Gradients for Discrete Random Variables by Sampling without ReplacementWouter Kool, Herke van Hoof, Max WellingICLR 2020 · 被引用 59 次
- Automatic Differentiation of Programs with Discrete RandomnessGaurav Arya, Moritz Schauer, Frank Schäfer, Christopher RackauckasNeurIPS 2022 · 被引用 56 次
- Differentiable Sampling of Categorical Distributions Using the CatLog-Derivative TrickLennert De Smet, Emanuele Sansone, Pedro Zuidberg Dos MartiresNeurIPS 2023 · 被引用 17 次
- Learning Randomly Perturbed Structured Predictors for Direct Loss MinimizationHedda Cohen Indelman, Tamir HazanICML 2021 · 被引用 9 次
它引用的顶会 Paper1
相关 Paper
- Local policy search with Bayesian optimizationSarah Müller, Alexander von Rohr, Sebastian TrimpeNeurIPS 2021 · 被引用 67 次
- Deterministic Value-Policy GradientsQingpeng Cai, Ling Pan, Pingzhong TangAAAI 2020 · 被引用 1 次
- From Importance Sampling to Doubly Robust Policy GradientJiawei Huang, Nan JiangICML 2020 · 被引用 26 次
- Guided Policy Optimization under Partial ObservabilityYueheng Li, Guangming Xie, Zongqing LuICLR 2026 · 被引用 4 次
- Learning to Schedule in Diffusion Probabilistic ModelsYunke Wang, Xiyu Wang, Anh-Dung Dinh, Bo Du 等KDD 2023 · 被引用 17 次
