Lune

NeurIPS2021Top-tier venue

The Pareto Frontier of model selection for general Contextual Bandits

Teodor Vanislavov Marinov, Julian Zimmert

2021Year
31Citations
9Top-tier citations

Abstract

Recent progress in model selection raises the question of the fundamental limits of these techniques. Under specific scrutiny has been model selection for general contextual bandits with nested policy classes, resulting in a COLT2020 open problem. It asks whether it is possible to obtain simultaneously the optimal single algorithm guarantees over all policies in a nested sequence of policy classes, or if otherwise this is possible for a trade-off α∈[12,1)\alpha\in[\frac{1}{2},1) between complexity term and time: ln⁡(∣Πm∣)1−αTα\ln(|\Pi_m|)^{1-\alpha}T^\alpha. We give a disappointing answer to this question. Even in the purely stochastic regime, the desired results are unobtainable. We present a Pareto frontier of up to logarithmic factors matching upper and lower bounds, thereby proving that an increase in the complexity term ln⁡(∣Πm∣)\ln(|\Pi_m|) independent of TT is unavoidable for general policy classes. As a side result, we also resolve a COLT2016 open problem concerning second-order bounds in full-information games.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ba0ab705-4a73-4a44-9223-35d14c482b4a

Cited by top-tier papers9

Ask how each one uses it

Builds on2

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines