Lune

NeurIPS2025顶会

Improved Best-of-Both-Worlds Regret for Bandits with Delayed Feedback

Ofir Schlisselberg, Tal Lancewicki, Peter Auer, Yishay Mansour

2025年份
2被引次数

摘要

We study the multi-armed bandit problem with adversarially chosen delays in the Best-of-Both-Worlds (BoBW) framework, which aims to achieve near-optimal performance in both stochastic and adversarial environments. While prior work has made progress toward this goal, existing algorithms suffer from significant gaps to the known lower bounds, especially in the stochastic settings. Our main contribution is a new algorithm that, up to logarithmic factors, matches the known lower bounds in each setting individually. In the adversarial case, our algorithm achieves regret of O~(KT+D)\widetilde{O}(\sqrt{KT} + \sqrt{D}), which is optimal up to logarithmic terms, where TT is the number of rounds, KK is the number of arms, and DD is the cumulative delay. In the stochastic case, we provide a regret bound which scale as ∑i:Δi>0(log⁡T/Δi)+1K∑Δiσmax\sum_{i:\Delta_i>0}\left(\log T/\Delta_i\right) + \frac{1}{K}\sum \Delta_i \sigma_{max}, where Δi\Delta_i is the sub-optimality gap of arm ii and σmax⁡\sigma_{\max} is the maximum number of missing observations. To the best of our knowledge, this is the first BoBW algorithm to simultaneously match the lower bounds in both stochastic and adversarial regimes in delayed environment. Moreover, even beyond the BoBW setting, our stochastic regret bound is the first to match the known lower bound under adversarial delays, improving the second term over the best known result by a factor of KK.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 88cde72e-8140-4d4d-b2a7-cfcccd7fa856

它引用的顶会 Paper11

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖