Lune

ICML2026顶会

Prior Diffusiveness and Regret in the Linear-Gaussian Bandit

Yifan Zhu, John Duchi, Benjamin Van Roy

2026年份
1被引次数

摘要

We prove that Thompson sampling exhibits O~(σdT+drTr(Σ0))\tilde{O}(\sigma d \sqrt{T} + d r \sqrt{\mathrm{Tr}(\Sigma_0)}) Bayesian regret in the linear-Gaussian bandit with a N(μ0,Σ0)\mathcal{N}(\mu_0, \Sigma_0) prior distribution on the coefficients, where dd is the dimension, TT is the time horizon, rr is the maximum ℓ2\ell_2 norm of the actions, and σ2\sigma^2 is the noise variance. In contrast to existing regret bounds, this shows that to within logarithmic factors, the prior-dependent ''burn-in'' term drTr(Σ0)d r \sqrt{\mathrm{Tr}(\Sigma_0)} decouples additively from the minimax (long run) regret d T. Previous regret bounds exhibit a multiplicative dependence on these terms. We establish these results via a new ''elliptical potential'' lemma, and also provide a lower bound indicating that the burn-in term is unavoidable.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext b525157a-bc23-4ed3-9b33-e63f9380478a

它引用的顶会 Paper1

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖