Lune

NeurIPS2023Top-tier venue

Anytime Model Selection in Linear Bandits

Parnian Kassraie, Nicolas Emmenegger, Andreas Krause, Aldo Pacchiano

2023Year
8Citations
3Top-tier citations

Abstract

Model selection in the context of bandit optimization is a challenging problem, as it requires balancing exploration and exploitation not only for action selection, but also for model selection. One natural approach is to rely on online learning algorithms that treat different models as experts. Existing methods, however, scale poorly (polyM\text{poly}M) with the number of models MM in terms of their regret. Our key insight is that, for model selection in linear bandits, we can emulate full-information feedback to the online learner with a favorable bias-variance trade-off. This allows us to develop ALEXP, which has an exponentially improved (log⁡M\log M) dependence on MM for its regret. ALEXP has anytime guarantees on its regret, and neither requires knowledge of the horizon nn, nor relies on an initial purely exploratory stage. Our approach utilizes a novel time-uniform analysis of the Lasso, establishing a new connection between online learning and high-dimensional statistics.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext ec20efa0-7d0e-4b60-ae67-7932ac8e7c4e

Cited by top-tier papers3

Ask how each one uses it

Builds on10

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines