Safe Linear Stochastic Bandits
Kia Khezeli, Eilyan Bitar
摘要
We introduce the safe linear stochastic bandit frameworka generalization of linear stochastic bandits-where, in each stage, the learner is required to select an arm with an expected reward that is no less than a predetermined (safe) threshold with high probability. We assume that the learner initially has knowledge of an arm that is known to be safe, but not necessarily optimal. Leveraging on this assumption, we introduce a learning algorithm that systematically combines known safe arms with exploratory arms to safely expand the set of safe arms over time, while facilitating safe greedy exploitation in subsequent stages. In addition to ensuring the satisfaction of the safety constraint at every stage of play, the proposed algorithm is shown to exhibit an expected regret that is no more than O( √ T log(T )) after T stages of play.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper15
- An Efficient Pessimistic-Optimistic Algorithm for Stochastic Linear Bandits with General ConstraintsXin Liu, Bin Li, Pengyi Shi, Lei YingNeurIPS 2021 · 被引用 63 次
- DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement LearningArchana Bura, Aria HasanzadeZonuzy, Dileep Kalathil, Srinivas Shakkottai 等NeurIPS 2022 · 被引用 48 次
- Stage-wise Conservative Linear BanditsAhmadreza Moradipari, Christos Thrampoulidis, Mahnoosh AlizadehNeurIPS 2020 · 被引用 37 次
- Safe Online Convex Optimization with Unknown Linear Safety ConstraintsSapana Chaudhary, Dileep M. KalathilAAAI 2022 · 被引用 21 次
- Interactively Learning Preference Constraints in Linear BanditsDavid Lindner, Sebastian Tschiatschek, Katja Hofmann, Andreas KrauseICML 2022 · 被引用 15 次
相关 Paper
- Feasible Action Search for Bandit Linear Programs via Thompson SamplingAditya Gangrade, Aldo Pacchiano, Clayton Scott, Venkatesh SaligramaICML 2025
- Preselection BanditsViktor Bengs, Eyke HüllermeierICML 2020 · 被引用 7 次
- Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and RiskTianrui Chen, Aditya Gangrade, Venkatesh SaligramaICML 2022 · 被引用 18 次
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
- Online Learning with Unknown ConstraintsKarthik Sridharan, Seung Won Wilson YooICML 2025
