Lune

ICML2025顶会

Triple-Optimistic Learning for Stochastic Contextual Bandits with General Constraints

Hengquan Guo, Lingkai Zu, Xin Liu

出版方
2025年份
2顶会引用

摘要

We study contextual bandits with general constraints, where a learner observes contexts and aims to maximize cumulative rewards while satisfying a wide range of general constraints. We introduce the Optimistic 3 framework, a novel learning and decision-making approach that integrates optimistic design into parameter learning, primal decision, and dual violation adaptation (i.e., triple-optimism), combined with an efficient primal-dual architecture. Optimistic 3 achieves Õ( √ T ) regret and constraint violation for contextual bandits with general constraints. This framework not only outperforms the stateof-the-art results that achieve Õ(T 3 4 ) guarantees when Slater's condition does not hold but also improves on previous results that achieve Õ( √ T /δ) when Slater's condition holds (δ denotes the Slater's condition parameter), offering a O(1/δ) improvement. Note this improvement is significant because δ can be arbitrarily small when constraints are particularly challenging. Moreover, we show that Optimistic 3 can be extended to classical multi-armed bandits with both stochastic and adversarial constraints, recovering the best-of-both-worlds guarantee established in the state-of-the-art works, but with significantly less computational overhead.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper12

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖