Strategies for Safe Multi-Armed Bandits with Logarithmic Regret and Risk
Tianrui Chen, Aditya Gangrade, Venkatesh Saligrama
摘要
We investigate a natural but surprisingly unstudied approach to the multi-armed bandit problem under safety risk constraints. Each arm is associated with an unknown law on safety risks and rewards, and the learner's goal is to maximise reward whilst not playing unsafe arms, as determined by a given threshold on the mean risk. We formulate a pseudo-regret for this setting that enforces this safety constraint in a per-round way by softly penalising any violation, regardless of the gain in reward due to the same. This has practical relevance to scenarios such as clinical trials, where one must maintain safety for each round rather than in an aggregated sense. We describe doubly optimistic strategies for this scenario, which maintain optimistic indices for both safety risk and reward. We show that schema based on both frequentist and Bayesian indices satisfy tight gap-dependent logarithmic regret bounds, and further that these play unsafe arms only logarithmically many times in total. This theoretical analysis is complemented by simulation studies demonstrating the effectiveness of the proposed schema, and probing the domains in which their use is appropriate.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Data-Dependent Regret Bounds for Constrained MABsGianmarco Genalti, Francesco Emanuele Stradi, Matteo Castiglioni, Alberto Marchesi 等NeurIPS 2025 · 被引用 3 次
- Triple-Optimistic Learning for Stochastic Contextual Bandits with General ConstraintsHengquan Guo, Lingkai Zu, Xin LiuICML 2025
- Pairwise Elimination with Instance-Dependent Guarantees for Bandits with Cost SubsidyIshank Juneja, Carlee Joe-Wong, Osman YaganICLR 2025
- Constrained Linear Thompson SamplingAditya Gangrade, Venkatesh SaligramaNeurIPS 2025
- Contextual Multi-Armed Bandits with Minimum Aggregated Revenue ConstraintsAhmed Ben Yahmed, Hafedh El Ferchichi, Marc Abeille, Vianney PerchetICLR 2026
相关 Paper
- Balancing Performance and Costs in Best Arm IdentificationMichael O. Harding, Kirthevasan KandasamyNeurIPS 2025 · 被引用 1 次
- Stage-wise Conservative Linear BanditsAhmadreza Moradipari, Christos Thrampoulidis, Mahnoosh AlizadehNeurIPS 2020 · 被引用 37 次
- Learning Policies with Zero or Bounded Constraint Violation for Constrained MDPsTao Liu, Ruida Zhou, Dileep Kalathil, Panganamala R. Kumar 等NeurIPS 2021 · 被引用 110 次
- DOPE: Doubly Optimistic and Pessimistic Exploration for Safe Reinforcement LearningArchana Bura, Aria HasanzadeZonuzy, Dileep Kalathil, Srinivas Shakkottai 等NeurIPS 2022 · 被引用 48 次
- Safe Linear Stochastic BanditsKia Khezeli, Eilyan BitarAAAI 2020 · 被引用 31 次
