Lune

ICML2026Top-tier venue

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

Yifan Zhu, John Duchi, Benjamin Van Roy

2026Year
1Citations

Abstract

We prove that Thompson sampling exhibits O~(σdT+drTr(Σ0))\tilde{O}(\sigma d \sqrt{T} + d r \sqrt{\mathrm{Tr}(\Sigma_0)}) Bayesian regret in the linear-Gaussian bandit with a N(μ0,Σ0)\mathcal{N}(\mu_0, \Sigma_0) prior distribution on the coefficients, where dd is the dimension, TT is the time horizon, rr is the maximum ℓ2\ell_2 norm of the actions, and σ2\sigma^2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ''burn-in'' term drTr(Σ0)d r \sqrt{\mathrm{Tr}(\Sigma_0)} decouples additively from the minimax (long run) regret d T. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ''elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext b525157a-bc23-4ed3-9b33-e63f9380478a

Builds on1

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines