Sleeping Reinforcement Learning
Simone Drago, Marco Mussi, Alberto Maria Metelli
摘要
In the standard Reinforcement Learning (RL) paradigm, the action space is assumed to be fixed and immutable throughout the learning process. However, in many real-world scenarios, not all actions are available at every decision stage. The available action set may depend on the current environment state, domain-specific constraints, or other (potentially stochastic) factors outside the agent's control. To address these realistic scenarios, we introduce a novel paradigm called Sleeping Reinforcement Learning, where the available action set varies during the interaction with the environment. We start with the simpler scenario in which the available action sets are revealed at the beginning of each episode. We show that a modification of UCBVI achieves regret of order r OpH ? SAT q, where H is the horizon, S and A are the cardinalities of the state and action spaces, respectively, and T is the learning horizon. Next, we address the more challenging and realistic scenario in which the available actions are disclosed only at each decision stage. By leveraging a novel construction, we establish a minimax lower bound of order Ωp ?
T 2 A2 q when the availability of actions is governed by a Markovian process, establishing a statistical barrier of the problem. Focusing on the statistically tractable case where action availability depends only on the current state and stage, we propose a new optimistic algorithm that achieves regret guarantees of order r OpH ? SAT q, showing that the problem shares the same complexity of standard RL.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Improved Sleeping Bandits with Stochastic Action Sets and Adversarial RewardsAadirupa Saha, Pierre Gaillard, Michal ValkoICML 2020 · 被引用 20 次
- Dueling Bandits with Adversarial SleepingAadirupa Saha, Pierre GaillardNeurIPS 2021 · 被引用 10 次
- Reinforcement Learning with Action-Triggered ObservationsAlexander Ryabchenko, Wenlong MouICML 2026 · 被引用 1 次
- Lifelong Learning with a Changing Action SetYash Chandak, Georgios Theocharous, Chris Nota, Philip S. ThomasAAAI 2020 · 被引用 39 次
- Achieving Õ(1/ε) Sample Complexity for Constrained Markov Decision ProcessJiashuo Jiang, Yinyu YeNeurIPS 2024 · 被引用 3 次
