ICML2026

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

Yifan Zhu, John Duchi, Benjamin Van Roy

被引用 1 次

摘要

We prove that Thompson sampling exhibits O~(σdT+drTr(Σ0))\tilde{O}(\sigma d \sqrt{T} + d r \sqrt{\mathrm{Tr}(\Sigma_0)}) Bayesian regret in the linear-Gaussian bandit with a N(μ0,Σ0)\mathcal{N}(\mu_0, \Sigma_0) prior distribution on the coefficients, where dd is the dimension, TT is the time horizon, rr is the maximum 2\ell_2 norm of the actions, and σ2\sigma^2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ''burn-in'' term drTr(Σ0)d r \sqrt{\mathrm{Tr}(\Sigma_0)} decouples additively from the minimax (long run) regret d T. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ''elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.