Safe Linear Stochastic Bandits
Kia Khezeli, Eilyan Bitar
Abstract
We introduce the safe linear stochastic bandit frameworka generalization of linear stochastic bandits-where, in each stage, the learner is required to select an arm with an expected reward that is no less than a predetermined (safe) threshold with high probability. We assume that the learner initially has knowledge of an arm that is known to be safe, but not necessarily optimal. Leveraging on this assumption, we introduce a learning algorithm that systematically combines known safe arms with exploratory arms to safely expand the set of safe arms over time, while facilitating safe greedy exploitation in subsequent stages. In addition to ensuring the satisfaction of the safety constraint at every stage of play, the proposed algorithm is shown to exhibit an expected regret that is no more than O( √ T log(T )) after T stages of play.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0ea3f578-0222-47ed-b84b-bdc9e325b102Cited by top-tier papers15
- An Efficient Pessimistic-Optimistic Algorithm for Stochastic Linear Bandits with General ConstraintsXin Liu, Bin Li, Pengyi Shi, Lei YingNeurIPS 2021 · 63 citations
- DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement LearningArchana Bura, Aria HasanzadeZonuzy, Dileep Kalathil, Srinivas Shakkottai et al.NeurIPS 2022 · 48 citations
- Stage-wise Conservative Linear BanditsAhmadreza Moradipari, Christos Thrampoulidis, Mahnoosh AlizadehNeurIPS 2020 · 37 citations
- Safe Online Convex Optimization with Unknown Linear Safety ConstraintsSapana Chaudhary, Dileep M. KalathilAAAI 2022 · 21 citations
- Interactively Learning Preference Constraints in Linear BanditsDavid Lindner, Sebastian Tschiatschek, Katja Hofmann, Andreas KrauseICML 2022 · 15 citations
Related papers
- Feasible Action Search for Bandit Linear Programs via Thompson SamplingAditya Gangrade, Aldo Pacchiano, Clayton Scott, Venkatesh SaligramaICML 2025
- Preselection BanditsViktor Bengs, Eyke HüllermeierICML 2020 · 7 citations
- Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and RiskTianrui Chen, Aditya Gangrade, Venkatesh SaligramaICML 2022 · 18 citations
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
- Online Learning with Unknown ConstraintsKarthik Sridharan, Seung Won Wilson YooICML 2025
