Lune

ICLR2021顶会

Impact of Representation Learning in Linear Bandits

Jiaqi Yang, Wei Hu, Jason D. Lee, Simon Shaolei Du

2021年份
58被引次数
21顶会引用

摘要

We study how representation learning can improve the efficiency of bandit problems. We study the setting where we play TT linear bandits with dimension dd concurrently, and these TT bandit tasks share a common k(≪d)k (\ll d) dimensional linear representation. For the finite-action setting, we present a new algorithm which achieves O~(TkN+dkNT)\widetilde{O}(T\sqrt{kN} + \sqrt{dkNT}) regret, where NN is the number of rounds we play for each bandit. When TT is sufficiently large, our algorithm significantly outperforms the naive algorithm (playing TT bandits independently) that achieves O~(TdN)\widetilde{O}(T\sqrt{d N}) regret. We also provide an Ω(TkN+dkNT)\Omega(T\sqrt{kN} + \sqrt{dkNT}) regret lower bound, showing that our algorithm is minimax-optimal up to poly-logarithmic factors. Furthermore, we extend our algorithm to the infinite-action setting and obtain a corresponding regret bound which demonstrates the benefit of representation learning in certain regimes. We also present experiments on synthetic and real-world data to illustrate our theoretical findings and demonstrate the effectiveness of our proposed algorithms.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext 669d69c9-f57e-4ee9-a89e-444d633bb4b4

引用它的顶会 Paper21

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖