An Online Learning Approach to Sequential User-Centric Selection Problems
Junpu Chen, Hong Xie
摘要
This paper proposes a new variant of multi-play MAB model, to capture important factors of the sequential user-centric selection problem arising from mobile edge computing, ridesharing applications, etc. In the proposed model, each arm is associated with discrete units of resources, each play is associate with movement costs and multiple plays can pull the same arm simultaneously. To learn the optimal action profile (an action profile prescribes the arm that each play pulls), there are two challenges: (1) the number of action profiles is large, i.e., M^K, where K and M denote the number of plays and arms respectively; (2) feedbacks on action profiles are not available, but instead feedbacks on some model parameters can be observed. To address the first challenge, we formulate a completed weighted bipartite graph to capture key factors of the offline decision problem with given model parameters. We identify the correspondence between action profiles and a special class of matchings of the graph. We also identify a dominance structure of this class of matchings. This correspondence and dominance structure enable us to design an algorithm named OffOptActPrf to locate the optimal action efficiently. To address the second challenge, we design an OnLinActPrf algorithm. We design estimators for model parameters and use these estimators to design a Quasi-UCB index for each action profile. The OnLinActPrf uses OffOptActPrf as a subroutine to select the action profile with the largest Quasi-UCB index. We conduct extensive experiments to validate the efficiency of OnLinActPrf.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Fractional Deep Reinforcement Learning for Age-Minimal Mobile Edge ComputingLyudong Jin, Ming Tang, Meng Zhang, Hao WangAAAI 2024 · 被引用 10 次
- Multiple-play Stochastic Bandits with Prioritized Arm Capacity SharingHong Xie, Haoran Gu, Yanying Huang, Tao Tan 等AAAI 2026
相关 Paper
- Multiple-Play Stochastic Bandits with Shareable Finite-Capacity ArmsXuchuang Wang, Hong Xie, John C. S. LuiICML 2022 · 被引用 8 次
- Socially-Optimal Mechanism Design for Incentivized Online LearningZhiyuan Wang, Lin Gao, Jianwei HuangINFOCOM 2022 · 被引用 11 次
- Online Restless Multi-Armed Bandits with Long-Term Fairness ConstraintsShufan Wang, Guojun Xiong, Jian LiAAAI 2024 · 被引用 11 次
- Bandit Learning with Joint Effect of Incentivized Sampling, Delayed Sampling Feedback, and Self-Reinforcing User PreferencesTianchen Zhou, Jia Liu, Chaosheng Dong, Yi SunICLR 2022 · 被引用 1 次
- Incentivized Bandit Learning with Self-Reinforcing User PreferencesTianchen Zhou, Jia Liu, Chaosheng Dong, Jingyuan DengICML 2021 · 被引用 2 次
