Lune

NeurIPS2023顶会

Blocked Collaborative Bandits: Online Collaborative Filtering with Per-Item Budget Constraints

Soumyabrata Pal, Arun Sai Suggala, Karthikeyan Shanmugam, Prateek Jain

2023年份
3被引次数

摘要

We consider the problem of blocked collaborative bandits where there are multiple users, each with an associated multi-armed bandit problem. These users are grouped into latent clusters such that the mean reward vectors of users within the same cluster are identical. Our goal is to design algorithms that maximize the cumulative reward accrued by all the users over time, under the constraint that no arm of a user is pulled more than B\mathsf{B} times. This problem has been originally considered by , and designing regret-optimal algorithms for it has since remained an open problem. In this work, we propose an algorithm called B-LATTICE (Blocked Latent bAndiTs via maTrIx ComplEtion) that collaborates across users, while simultaneously satisfying the budget constraints, to maximize their cumulative rewards. Theoretically, under certain reasonable assumptions on the latent structure, with M\mathsf{M} users, N\mathsf{N} arms, T\mathsf{T} rounds per user, and C=O(1)\mathsf{C}=O(1) latent clusters, B-LATTICE achieves a per-user regret of O~(T(1+NM−1)\widetilde{O}(\sqrt{\mathsf{T}(1 + \mathsf{N}\mathsf{M}^{-1})} under a budget constraint of B=Θ(log⁡T)\mathsf{B}=\Theta(\log \mathsf{T}). These are the first sub-linear regret bounds for this problem, and match the minimax regret bounds when B=T\mathsf{B}=\mathsf{T}. Empirically, we demonstrate that our algorithm has superior performance over baselines even when B=1\mathsf{B}=1. B-LATTICE runs in phases where in each phase it clusters users into groups and collaborates across users within a group to quickly learn their reward models.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖