Lune

ICML2022顶会

Feature and Parameter Selection in Stochastic Linear Bandits

Ahmadreza Moradipari, Berkay Turan, Yasin Abbasi-Yadkori, Mahnoosh Alizadeh, Mohammad Ghavamzadeh

2022年份
6被引次数
2顶会引用

摘要

We study two model selection settings in stochastic linear bandits (LB). In the first setting, which we refer to as feature selection, the expected reward of the LB problem is in the linear span of at least one of MM feature maps (models). In the second setting, the reward parameter of the LB problem is arbitrarily selected from MM models represented as (possibly) overlapping balls in Rd\mathbb R^d. However, the agent only has access to misspecified models, i.e., estimates of the centers and radii of the balls. We refer to this setting as parameter selection. For each setting, we develop and analyze a computationally efficient algorithm that is based on a reduction from bandits to full-information problems. This allows us to obtain regret bounds that are not worse (up to a log⁡M\sqrt{\log M} factor) than the case where the true model is known. This is the best-reported dependence on the number of models MM in these settings. Finally, we empirically show the effectiveness of our algorithms using synthetic and real-world experiments.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper2

问问它们各自怎么用它

它引用的顶会 Paper8

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖