Reinforcement Learning When All Actions Are Not Always Available
Yash Chandak, Georgios Theocharous, Blossom Metevier, Philip S. Thomas
摘要
The Markov decision process (MDP) formulation used to model many real-world sequential decision making problems does not efficiently capture the setting where the set of available decisions (actions) at each time step is stochastic. Recently, the stochastic action set Markov decision process (SAS-MDP) formulation has been proposed, which better captures the concept of a stochastic action set. In this paper we argue that existing RL algorithms for SAS-MDPs can suffer from potential divergence issues, and present new policy gradient algorithms for SAS-MDPs that incorporate variance reduction techniques unique to this setting, and provide conditions for their convergence. We conclude with experiments that demonstrate the practicality of our approaches on tasks inspired by real-life use cases wherein the action set is stochastic. Potential Limitations of SAS-Q-Learning Although SAS-Q-learning provides a powerful first modelfree algorithm for approximating optimal policies for SAS-MDPs, it inherits several of the drawbacks of the Q-learning
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- On the Convergence Theory of Debiased Model-Agnostic Meta-Reinforcement LearningAlireza Fallah, Kristian Georgiev, Aryan Mokhtari, Asuman E. OzdaglarNeurIPS 2021 · 被引用 31 次
- Model-Free Robust Average-Reward Reinforcement LearningYue Wang, Alvaro Velasquez, George K. Atia, Ashley Prater-Bennette 等ICML 2023 · 被引用 25 次
- Acting in Delayed Environments with Non-Stationary Markov PoliciesEsther Derman, Gal Dalal, Shie MannorICLR 2021 · 被引用 4 次
- Reinforcement Learning with Non-Markovian RewardsMaor Gaon, Ronen I. BrafmanAAAI 2020 · 被引用 96 次
- A Single-Loop Robust Policy Gradient Method for Robust Markov Decision ProcessesZhenwei Lin, Chenyu Xue, Qi Deng, Yinyu YeICML 2024 · 被引用 3 次
