Lune

NeurIPS2022顶会

Contextual Bandits with Knapsacks for a Conversion Model

Zhen Li, Gilles Stoltz

2022年份
4被引次数
5顶会引用

摘要

We consider contextual bandits with knapsacks, with an underlying structure between rewards generated and cost vectors suffered. We do so motivated by sales with commercial discounts. At each round, given the stochastic i.i.d. context xt\mathbf{x}_t and the arm picked ata_t (corresponding, e.g., to a discount level), a customer conversion may be obtained, in which case a reward r(a,xt)r(a,\mathbf{x}_t) is gained and vector costs c(at,xt)c(a_t,\mathbf{x}_t) are suffered (corresponding, e.g., to losses of earnings). Otherwise, in the absence of a conversion, the reward and costs are null. The reward and costs achieved are thus coupled through the binary variable measuring conversion or the absence thereof. This underlying structure between rewards and costs is different from the linear structures considered by Agrawal and Devanur [2016] (but we show that the techniques introduced in the present article may also be applied to the case of these linear structures). The adaptive policies exhibited solve at each round a linear program based on upper-confidence estimates of the probabilities of conversion given aa and x\mathbf{x}. This kind of policy is most natural and achieves a regret bound of the typical order (OPT/BB) T\sqrt{T}, where BB is the total budget allowed, OPT is the optimal expected reward achievable by a static policy, and TT is the number of rounds.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

lune papers fulltext aa902dfb-f972-4fdb-a7f2-4951e21929c3

引用它的顶会 Paper5

问问它们各自怎么用它

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖