Towards Safe Policy Improvement for Non-Stationary MDPs
Yash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White, Philip S. Thomas
摘要
Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that are safe for deployment, they assume that the underlying problem is stationary. However, many real-world problems of interest exhibit non-stationarity, and when stakes are high, the cost associated with a false stationarity assumption may be unacceptable. We take the first steps towards ensuring safety, with high confidence, for smoothly-varying non-stationary decision problems. Our proposed method extends a type of safe algorithm, called a Seldonian algorithm, through a synthesis of model-free reinforcement learning with time-series analysis. Safety is ensured using sequential hypothesis testing of a policy's forecasted performance, and confidence intervals are obtained using wild bootstrap.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper8
- An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextXiaoyu Chen, Xiangming Zhu, Yufeng Zheng, Pushi Zhang 等NeurIPS 2022 · 被引用 24 次
- Continual Auxiliary Task LearningMatthew McLeod, Chunlok Lo, Matthew Schlegel, Andrew Jacobsen 等NeurIPS 2021 · 被引用 13 次
- Scalable Safe Policy Improvement via Monte Carlo Tree SearchAlberto Castellini, Federico Bianchi, Edoardo Zorzi, Thiago D. Simão 等ICML 2023 · 被引用 9 次
- Off-Policy Evaluation for Action-Dependent Non-stationary EnvironmentsYash Chandak, Shiv Shankar, Nathaniel D. Bastian, Bruno C. da Silva 等NeurIPS 2022 · 被引用 7 次
- Tempo Adaptation in Non-stationary Reinforcement LearningHyunin Lee, Yuhao Ding, Jongmin Lee, Ming Jin 等NeurIPS 2023 · 被引用 6 次
它引用的顶会 Paper4
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 被引用 285 次
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White 等ICML 2020 · 被引用 72 次
- Lipschitz Lifelong Reinforcement LearningErwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai 等AAAI 2021 · 被引用 43 次
- Lifelong Learning with a Changing Action SetYash Chandak, Georgios Theocharous, Chris Nota, Philip S. ThomasAAAI 2020 · 被引用 39 次
相关 Paper
- High Confidence Generalization for Reinforcement LearningJames E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous 等ICML 2021 · 被引用 5 次
- Security Analysis of Safe and Seldonian Reinforcement Learning AlgorithmsPinar Ozisik, Philip S. ThomasNeurIPS 2020 · 被引用 4 次
- SafeDICE: Offline Safe Imitation Learning with Non-Preferred DemonstrationsYoungsoo Jang, Geon-Hyeong Kim, Jongmin Lee, Sungryull Sohn 等NeurIPS 2023 · 被引用 9 次
- Safe Time-Varying Optimization based on Gaussian Processes with Spatio-Temporal KernelJialin Li, Marta Zagórowska, Giulia De Pasquale, Alisa Rupenyan 等NeurIPS 2024 · 被引用 11 次
- Near-Optimal Model-Free Reinforcement Learning in Non-Stationary Episodic MDPsWeichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi 等ICML 2021 · 被引用 49 次
