Socially-Optimal Mechanism Design for Incentivized Online Learning
Zhiyuan Wang, Lin Gao, Jianwei Huang
Abstract
Multi-arm bandit (MAB) is a classic online learning framework that studies the sequential decision-making in an uncertain environment. The MAB framework, however, overlooks the scenario where the decision-maker cannot take actions (e.g., pulling arms) directly. It is a practically important scenario in many applications such as spectrum sharing, crowdsensing, and edge computing. In these applications, the decision-maker would incentivize other selfish agents to carry out desired actions (i.e., pulling arms on the decision-maker’s behalf). This paper establishes the incentivized online learning (IOL) framework for this scenario. The key challenge to design the IOL framework lies in the tight coupling of the unknown environment learning and asymmetric information revelation. To address this, we construct a special Lagrangian function based on which we propose a socially-optimal mechanism for the IOL framework. Our mechanism satisfies various desirable properties such as agent fairness, incentive compatibility, and voluntary participation. It achieves the same asymptotic performance as the state-of-art benchmark that requires extra information. Our analysis also unveils the power of crowd in the IOL framework: a larger agent crowd enables our mechanism to approach more closely the theoretical upper bound of social performance. Numerical results demonstrate the advantages of our mechanism in large-scale edge computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c34049ac-a7c3-4c6a-a702-1ba3ccec54e9Cited by top-tier papers1
Ask how each one uses itBuilds on1
Related papers
- Multi-Agent Distributed Reinforcement Learning for Making Decentralized Offloading DecisionsJing Tan, Ramin Khalili, Holger Karl, Artur HeckerINFOCOM 2022 · 33 citations
- Near Optimal and Dynamic Mechanisms Towards a Stable NFV Market in Multi-Tier Cloud NetworksZichuan Xu, Haozhe Ren, Weifa Liang, Qiufen Xia et al.INFOCOM 2021 · 11 citations
- Marginal Value-Based Edge Resource Pricing and Allocation for Deadline-Sensitive TasksPuwei Wang, Zhouxing Sun, Ying Zhan, Haoran Li et al.INFOCOM 2023 · 8 citations
- An Incentive Mechanism Design for Efficient Edge Learning by Deep Reinforcement Learning ApproachYufeng Zhan, Jiang ZhangINFOCOM 2020 · 102 citations
- Decentralized Task Offloading in Edge Computing: A Multi-User Multi-Armed Bandit ApproachXiong Wang, Jiancheng Ye, John C. S. LuiINFOCOM 2022 · 89 citations
