Towards Safe Policy Improvement for Non-Stationary MDPs
Yash Chandak, Scott M. Jordan, Georgios Theocharous, Martha White, Philip S. Thomas
Abstract
Many real-world sequential decision-making problems involve critical systems with financial risks and human-life risks. While several works in the past have proposed methods that are safe for deployment, they assume that the underlying problem is stationary. However, many real-world problems of interest exhibit non-stationarity, and when stakes are high, the cost associated with a false stationarity assumption may be unacceptable. We take the first steps towards ensuring safety, with high confidence, for smoothly-varying non-stationary decision problems. Our proposed method extends a type of safe algorithm, called a Seldonian algorithm, through a synthesis of model-free reinforcement learning with time-series analysis. Safety is ensured using sequential hypothesis testing of a policy's forecasted performance, and confidence intervals are obtained using wild bootstrap.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext af2bfd06-5a6d-4c23-9159-e8ce8f5e550dCited by top-tier papers8
- An Adaptive Deep RL Method for Non-Stationary Environments with Piecewise Stable ContextXiaoyu Chen, Xiangming Zhu, Yufeng Zheng, Pushi Zhang et al.NeurIPS 2022 · 24 citations
- Continual Auxiliary Task LearningMatthew McLeod, Chunlok Lo, Matthew Schlegel, Andrew Jacobsen et al.NeurIPS 2021 · 13 citations
- Scalable Safe Policy Improvement via Monte Carlo Tree SearchAlberto Castellini, Federico Bianchi, Edoardo Zorzi, Thiago D. Simão et al.ICML 2023 · 9 citations
- Off-Policy Evaluation for Action-Dependent Non-stationary EnvironmentsYash Chandak, Shiv Shankar, Nathaniel D. Bastian, Bruno C. da Silva et al.NeurIPS 2022 · 7 citations
- Tempo Adaptation in Non-stationary Reinforcement LearningHyunin Lee, Yuhao Ding, Jongmin Lee, Ming Jin et al.NeurIPS 2023 · 6 citations
Builds on4
- Fast Differentiable Sorting and RankingMathieu Blondel, Olivier Teboul, Quentin Berthet, Josip DjolongaICML 2020 · 285 citations
- Optimizing for the Future in Non-Stationary MDPsYash Chandak, Georgios Theocharous, Shiv Shankar, Martha White et al.ICML 2020 · 72 citations
- Lipschitz Lifelong Reinforcement LearningErwan Lecarpentier, David Abel, Kavosh Asadi, Yuu Jinnai et al.AAAI 2021 · 43 citations
- Lifelong Learning with a Changing Action SetYash Chandak, Georgios Theocharous, Chris Nota, Philip S. ThomasAAAI 2020 · 39 citations
Related papers
- High Confidence Generalization for Reinforcement LearningJames E. Kostas, Yash Chandak, Scott M. Jordan, Georgios Theocharous et al.ICML 2021 · 5 citations
- Security Analysis of Safe and Seldonian Reinforcement Learning AlgorithmsPinar Ozisik, Philip S. ThomasNeurIPS 2020 · 4 citations
- SafeDICE: Offline Safe Imitation Learning with Non-Preferred DemonstrationsYoungsoo Jang, Geon-Hyeong Kim, Jongmin Lee, Sungryull Sohn et al.NeurIPS 2023 · 9 citations
- Safe Time-Varying Optimization based on Gaussian Processes with Spatio-Temporal KernelJialin Li, Marta Zagórowska, Giulia De Pasquale, Alisa Rupenyan et al.NeurIPS 2024 · 11 citations
- Near-Optimal Model-Free Reinforcement Learning in Non-Stationary Episodic MDPsWeichao Mao, Kaiqing Zhang, Ruihao Zhu, David Simchi-Levi et al.ICML 2021 · 49 citations
