Multi-Agent Reinforcement Learning in Stochastic Networked Systems
Yiheng Lin, Guannan Qu, Longbo Huang, Adam Wierman
摘要
We study multi-agent reinforcement learning (MARL) in a stochastic network of agents. The objective is to find localized policies that maximize the (discounted) global reward. In general, scalability is a challenge in this setting because the size of the global state/action space can be exponential in the number of agents. Scalable algorithms are only known in cases where dependencies are static, fixed and local, e.g., between neighbors in a fixed, time-invariant underlying graph. In this work, we propose a Scalable Actor Critic framework that applies in settings where the dependencies can be non-local and stochastic, and provide a finite-time error bound that shows how the convergence rate depends on the speed of information spread in the network. Additionally, as a byproduct of our analysis, we obtain novel finite-time convergence results for a general stochastic approximation scheme and for temporal difference learning with state aggregation, which apply beyond the setting of MARL in networked systems.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Scalable Primal-Dual Actor-Critic Method for Safe Multi-Agent RL with General UtilitiesDonghao Ying, Yunkai Zhang, Yuhao Ding, Alec Koppel 等NeurIPS 2023 · 被引用 28 次
- Scalable Constrained Policy Optimization for Safe Multi-agent Reinforcement LearningLijun Zhang, Lin Li, Wei Wei, Huizhong Song 等NeurIPS 2024 · 被引用 22 次
- Bayesian Ego-graph Inference for Networked Multi-Agent Reinforcement LearningWei Duan, Jie Lu, Junyu XuanNeurIPS 2025 · 被引用 15 次
- Locally Interdependent Multi-Agent MDP: Theoretical Framework for Decentralized Agents with Dynamic DependenciesAlex DeWeese, Guannan QuICML 2024 · 被引用 6 次
- Causality Meets Locality: Provably Generalizable and Scalable Policy Learning for Networked SystemsHao Liang, Shuqing Shi, Yudi Zhang, Biwei Huang 等NeurIPS 2025 · 被引用 1 次
它引用的顶会 Paper3
- A Finite-Time Analysis of Two Time-Scale Actor-Critic MethodsYue Wu, Weitong Zhang, Pan Xu, Quanquan GuNeurIPS 2020 · 被引用 189 次
- Q-learning with UCB Exploration is Sample Efficient for Infinite-Horizon MDPYuanhao Wang, Kefan Dong, Xiaoyu Chen, Liwei WangICLR 2020 · 被引用 107 次
- Scalable Multi-Agent Reinforcement Learning for Networked Systems with Average RewardGuannan Qu, Yiheng Lin, Adam Wierman, Na LiNeurIPS 2020 · 被引用 99 次
相关 Paper
- Local Policies for Graph-Structured Markov Decision ProcessesFathima Faizal, Asuman Ozdaglar, Martin WainwrightICML 2026
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement LearningZhiyao Zhang, Myeung Suk Oh, Hairi, Ziyue Luo 等ICML 2025
- A Law of Iterated Logarithm for Multi-Agent Reinforcement LearningGugan Thoppe, Bhumesh KumarNeurIPS 2021 · 被引用 4 次
- Decentralized TD Tracking with Linear Function Approximation and its Finite-Time AnalysisGang Wang, Songtao Lu, Georgios B. Giannakis, Gerald Tesauro 等NeurIPS 2020 · 被引用 30 次
- Scalable Multi-Agent Reinforcement Learning through Intelligent Information AggregationSiddharth Nayak, Kenneth Choi, Wenqi Ding, Sydney Dolan 等ICML 2023 · 被引用 73 次
