Lune

ICML2026顶会

Stochastic Linear Bandits with Parameter Noise

Daniel Ezer, Alon Peled-Cohen, Yishay Mansour

2026年份

摘要

We study the stochastic linear bandits with parameter noise model, in which the reward of action aa is a⊤θa^\top \theta where θ\theta is sampled i.i.d. We show a regret upper bound of O~(dTlog⁡(K/δ)σmax⁡2)\widetilde{O} (\sqrt{d T \log(K/\delta) \sigma^2_{\max}}) for a horizon TT, general action set of size KK of dimension dd, and where σmax⁡2\sigma^2_{\max} is the maximal variance of the reward for any action. We further provide a lower bound of Ω~(dTσmax⁡2)\widetilde{\Omega} (d \sqrt{T \sigma_{\max}^2}) which is tight (up to logarithmic factors) whenever log⁡K≈d\log K \approx d. For more specific action sets, ℓp\ell_p unit balls with p≤2p \leq 2 and dual norm qq, we show that the minimax regret is Θ~(dTσq2)\widetilde{\Theta} (\sqrt{dT \sigma_q^2}), where σq2\sigma_q^2 is a variance-dependent quantity that is always at most 44. This is in contrast to the minimax regret attainable for such sets in the classic additive noise model where the regret is of order dTd \sqrt{T}. Surprisingly, we show that this optimal (up to logarithmic factors) regret bound is attainable using a very simple explore-exploit algorithm.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper5

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖