Making the most of your day: online learning for optimal allocation of time
Etienne Boursier, Tristan Garrec, Vianney Perchet, Marco Scarsini
摘要
We study online learning for optimal allocation when the resource to be allocated is time. %Examples of possible applications include job scheduling for a computing server, a driver filling a day with rides, a landlord renting an estate, etc. An agent receives task proposals sequentially according to a Poisson process and can either accept or reject a proposed task. If she accepts the proposal, she is busy for the duration of the task and obtains a reward that depends on the task duration. If she rejects it, she remains on hold until a new task proposal arrives. We study the regret incurred by the agent, first when she knows her reward function but does not know the distribution of the task duration, and then when she does not know her reward function, either. This natural setting bears similarities with contextual (one-armed) bandits, but with the crucial difference that the normalized reward associated to a context depends on the whole distribution of contexts.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
相关 Paper
- Online Learning and Pricing for Network Revenue Management with Reusable ResourcesHuiwen Jia, Cong Shi, Siqian ShenNeurIPS 2022 · 被引用 8 次
- Online Learning and Pricing with Reusable Resources: Linear Bandits with Sub-Exponential RewardsHuiwen Jia, Cong Shi, Siqian ShenICML 2022 · 被引用 9 次
- Online Resource Allocation with Non-Stationary CustomersXiaoyue Zhang, Hanzhang Qin, Mabel C. ChouICML 2024
- Online Task Assignment Problems with Reusable ResourcesHanna Sumita, Shinji Ito, Kei Takemura, Daisuke Hatano 等AAAI 2022 · 被引用 10 次
- Learning to Schedule Tasks with Deadline and Throughput ConstraintsQingsong Liu, Zhixuan FangINFOCOM 2023 · 被引用 19 次
