Lune

ICML2021顶会

Problem Dependent View on Structured Thresholding Bandit Problems

James Cheshire, Pierre Ménard, Alexandra Carpentier

2021年份
8被引次数
2顶会引用

摘要

We investigate the problem dependent regime in the stochastic Thresholding Bandit problem (TBP) under several shape constraints. In the TBP, the objective of the learner is to output, at the end of a sequential game, the set of arms whose means are above a given threshold. The vanilla, unstructured, case is already well studied in the literature. Taking KK as the number of arms, we consider the case where (i) the sequence of arm's means (μk)k=1K(\mu_k)_{k=1}^K is monotonically increasing (MTBP) and (ii) the case where (μk)k=1K(\mu_k)_{k=1}^K is concave (CTBP). We consider both cases in the problem dependent regime and study the probability of error - i.e. the probability to mis-classify at least one arm. In the fixed budget setting, we provide upper and lower bounds for the probability of error in both the concave and monotone settings, as well as associated algorithms. In both settings the bounds match in the problem dependent regime up to universal constants in the exponential.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖