Lune

NeurIPS2021顶会

Contextual Recommendations and Low-Regret Cutting-Plane Algorithms

Sreenivas Gollapudi, Guru Guruganesh, Kostas Kollias, Pasin Manurangsi, Renato Paes Leme, Jon Schneider

2021年份
17被引次数
4顶会引用

摘要

We consider the following variant of contextual linear bandits motivated by routing applications in navigational engines and recommendation systems. We wish to learn a hidden d -dimensional value w ∗ . Every round, we are presented with a subset X t ⊆ R d of possible actions. If we choose (i.e. recommend to the user) action x t , we obtain utility (cid:104) x t , w ∗ (cid:105) but only learn the identity of the best action arg max x ∈X t (cid:104) x, w ∗ (cid:105) . We design algorithms for this problem which achieve regret O ( d log T ) and exp( O ( d log d )) . To accomplish this, we design novel cutting-plane algorithms with low “regret” – the total distance between the true point w ∗ and the hyperplanes the separation oracle returns. We also consider the variant where we are allowed to provide a list of several recommendations. In this variant, we give an algorithm with O ( d 2 log d ) regret and list size poly( d ) . Finally, we construct nearly tight algorithms for a weaker variant of this problem where the learner only learns the identity of an action that is better than the recommendation. Our results rely on new algorithmic techniques in convex geometry (including a variant of Steiner’s formula for the centroid of a convex set) which may be of independent interest.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

引用它的顶会 Paper4

问问它们各自怎么用它

它引用的顶会 Paper4

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖