Projection-based Lyapunov method for fully heterogeneous weakly-coupled MDPs
Xiangcheng Zhang, Yige Hong, Weina Wang
摘要
Heterogeneity poses a fundamental challenge for many real-world large-scale decision-making problems but remains largely understudied. In this paper, we study the fully heterogeneous setting of a prominent class of such problems, known as weakly-coupled Markov decision processes (WCMDPs). Each WCMDP consists of arms (or subproblems), which have distinct model parameters in the fully heterogeneous setting, leading to the curse of dimensionality when is large. We show that, under mild assumptions, an efficiently computable policy achieves an optimality gap in the long-run average reward per arm for fully heterogeneous WCMDPs as becomes large. This is the first asymptotic optimality result for fully heterogeneous average-reward WCMDPs. Our main technical innovation is the construction of projection-based Lyapunov functions that certify the convergence of rewards and costs to an optimal region, even under full heterogeneity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper8
- NeurWIN: Neural Whittle Index Network For Restless Bandits Via Deep RLKhaled Nakhleh, Santosh Ganji, Ping-Chun Hsieh, I-Hong Hou 等NeurIPS 2021 · 被引用 52 次
- Finite-Time Analysis of Whittle Index based Q-Learning for Restless Multi-Armed Bandits with Neural Network Function ApproximationGuojun Xiong, Jian LiNeurIPS 2023 · 被引用 23 次
- Learning Infinite-Horizon Average-Reward Restless Multi-Action Bandits via Index AwarenessGuojun Xiong, Shufan Wang, Jian LiNeurIPS 2022 · 被引用 21 次
- Restless Bandits with Average Reward: Breaking the Uniform Global Attractor AssumptionYige Hong, Qiaomin Xie, Yudong Chen, Weina WangNeurIPS 2023 · 被引用 13 次
- Q-Learning Lagrange Policies for Multi-Action Restless BanditsJackson A. Killian, Arpita Biswas, Sanket Shah, Milind TambeKDD 2021 · 被引用 12 次
相关 Paper
- Natural Policy Gradient Primal-Dual Method for Constrained Markov Decision ProcessesDongsheng Ding, Kaiqing Zhang, Tamer Basar, Mihailo R. JovanovicNeurIPS 2020 · 被引用 252 次
- Weakly Coupled Deep Q-NetworksIbrahim El Shar, Daniel R. JiangNeurIPS 2023 · 被引用 12 次
- Optimal Common Contract with Heterogeneous AgentsShenke Xiao, Zihe Wang, Mengjing Chen, Pingzhong Tang 等AAAI 2020 · 被引用 14 次
- Learning Human-Robot Collaboration via Heterogeneous-Agent Lyapunov Policy OptimizationHao Zhang, Yaru Niu, Yikai Wang, Ding Zhao 等ICML 2026
- Policy Optimization for Robust Average Reward MDPsZhongchang Sun, Sihong He, Fei Miao, Shaofeng ZouNeurIPS 2024 · 被引用 10 次
