Tilted Quantile Gradient Updates for Quantile-Constrained Reinforcement Learning
Chenglin Li, Guangchun Ruan, Hua Geng
摘要
Safe reinforcement learning (RL) is a popular and versatile paradigm to learn reward-maximizing policies with safety guarantees. Previous works tend to express the safety constraints in an expectation form due to the ease of implementation, but this turns out to be ineffective in maintaining safety constraints with high probability. To this end, we move to the quantile-constrained RL that enables a higher level of safety without any expectation-form approximations. We directly estimate the quantile gradients through sampling and provide the theoretical proofs of convergence. Then a tilted update strategy for quantile gradients is implemented to compensate the asymmetric distributional density, with a direct benefit of return performance. Experiments demonstrate that the proposed model fully meets safety requirements (quantile constraints) while outperforming the state-of-the-art benchmarks with higher return.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper3
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
- IPO: Interior-Point Policy Optimization under ConstraintsYongshuai Liu, Jiaxin Ding, Xin LiuAAAI 2020 · 被引用 231 次
- Quantile Constrained Reinforcement Learning: A Reinforcement Learning Framework Constraining Outage ProbabilityWhiyoung Jung, Myungsik Cho, Jongeui Park, Youngchul SungNeurIPS 2022 · 被引用 12 次
相关 Paper
- Constrained Variational Policy Optimization for Safe Reinforcement LearningZuxin Liu, Zhepeng Cen, Vladislav Isenbaev, Wei Liu 等ICML 2022 · 被引用 112 次
- Off-Policy Safe Reinforcement Learning with Cost-Constrained Optimistic ExplorationGuopeng Li, Matthijs T. J. Spaan, Julian F. P. KooijICLR 2026
- Density Constrained Reinforcement LearningZengyi Qin, Yuxiao Chen, Chuchu FanICML 2021 · 被引用 40 次
- Distributional Reinforcement Learning with Monotonic SplinesYudong Luo, Guiliang Liu, Haonan Duan, Oliver Schulte 等ICLR 2022 · 被引用 18 次
- Constraints Penalized Q-learning for Safe Offline Reinforcement LearningHaoran Xu, Xianyuan Zhan, Xiangyu ZhuAAAI 2022 · 被引用 127 次
