On the Hidden Biases of Policy Mirror Ascent in Continuous Action Spaces
Amrit Singh Bedi, Souradip Chakraborty, Anjaly Parayil, Brian M. Sadler, Pratap Tokekar, Alec Koppel
摘要
We focus on parameterized policy search for reinforcement learning over continuous action spaces. Typically, one assumes the score function associated with a policy is bounded, which fails to hold even for Gaussian policies. To properly address this issue, one must introduce an exploration tolerance parameter to quantify the region in which it is bounded. Doing so incurs a persistent bias that appears in the attenuation rate of the expected policy gradient norm, which is inversely proportional to the radius of the action space. To mitigate this hidden bias, heavy-tailed policy parameterizations may be used, which exhibit a bounded score function, but doing so can cause instability in algorithmic updates. To address these issues, in this work, we study the convergence of policy gradient algorithms under heavy-tailed parameterizations, which we propose to stabilize with a combination of mirror ascent-type updates and gradient tracking. Our main theoretical contribution is the establishment that this scheme converges with constant step and batch sizes, whereas prior works require these parameters to respectively shrink to null or grow to infinity. Experimentally, this scheme under a heavy-tailed policy parameterization yields improved reward accumulation across a variety of settings as compared with standard benchmarks.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- PARL: A Unified Framework for Policy Alignment in Reinforcement Learning from Human FeedbackSouradip Chakraborty, Amrit Singh Bedi, Alec Koppel, Huazheng Wang 等ICLR 2024 · 被引用 42 次
- Beyond Exponentially Fast Mixing in Average-Reward Reinforcement Learning via Multi-Level Monte Carlo Actor-CriticWesley A. Suttle, Amrit S. Bedi, Bhrij Patel, Brian M. Sadler 等ICML 2023 · 被引用 24 次
- Reinforcement Learning with General Utilities: Simpler Variance Reduction and Large State-Action SpaceAnas Barakat, Ilyas Fatkhullin, Niao HeICML 2023 · 被引用 18 次
- STEERING : Stein Information Directed Exploration for Model-Based Reinforcement LearningSouradip Chakraborty, Amrit S. Bedi, Alec Koppel, Mengdi Wang 等ICML 2023 · 被引用 9 次
- Towards Global Optimality for Practical Average Reward Reinforcement Learning without Mixing Time OraclesBhrij Patel, Wesley A. Suttle, Alec Koppel, Vaneet Aggarwal 等ICML 2024 · 被引用 4 次
它引用的顶会 Paper7
- Sample Efficient Reinforcement Learning with REINFORCEJunzi Zhang, Jongho Kim, Brendan O'Donoghue, Stephen P. BoydAAAI 2021 · 被引用 162 次
- Mirror Descent Policy OptimizationManan Tomar, Lior Shani, Yonathan Efroni, Mohammad GhavamzadehICLR 2022 · 被引用 111 次
- Sample Efficient Policy Gradient Methods with Recursive Variance ReductionPan Xu, Felicia Gao, Quanquan GuICLR 2020 · 被引用 99 次
- Hausdorff Dimension, Heavy Tails, and Generalization in Neural NetworksUmut Simsekli, Ozan Sener, George Deligiannidis, Murat A. ErdogduNeurIPS 2020 · 被引用 79 次
- Fractional Underdamped Langevin Dynamics: Retargeting SGD with Momentum under Heavy-Tailed Gradient NoiseUmut Simsekli, Lingjiong Zhu, Yee Whye Teh, Mert GürbüzbalabanICML 2020 · 被引用 58 次
相关 Paper
- Truncated Gaussian Policy for Debiased Continuous ControlGanghun Lee, Minji Kim, Minsu Lee, Byoung-Tak ZhangAAAI 2025 · 被引用 1 次
- Provably Robust Temporal Difference Learning for Heavy-Tailed RewardsSemih Cayci, Atilla EryilmazNeurIPS 2023 · 被引用 12 次
- q-exponential family for policy optimizationLingwei Zhu, Haseeb Shah, Han Wang, Yukie Nagai 等ICLR 2025
- Catastrophic Goodhart: regularizing RLHF with KL divergence does not mitigate heavy-tailed reward misspecificationThomas Kwa, Drake Thomas, Adrià Garriga-AlonsoNeurIPS 2024 · 被引用 22 次
- Exact Policy Recovery in Offline RL with Both Heavy-Tailed Rewards and Data CorruptionYiding Chen, Xuezhou Zhang, Qiaomin Xie, Xiaojin ZhuAAAI 2024 · 被引用 2 次
