Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement Learning
Tong Yang, Shicong Cen, Yuting Wei, Yuxin Chen, Yuejie Chi
摘要
Federated reinforcement learning (RL) enables collaborative decision making of multiple distributed agents without sharing local data trajectories. In this work, we consider a multi-task setting, in which each agent has its own private reward function corresponding to different tasks, while sharing the same transition kernel of the environment. Focusing on infinite-horizon Markov decision processes, the goal is to learn a globally optimal policy that maximizes the sum of the discounted total rewards of all the agents in a decentralized manner, where each agent only communicates with its neighbors over some prescribed graph topology. We develop federated vanilla and entropy-regularized natural policy gradient (NPG) methods in the tabular setting under softmax parameterization, where gradient tracking is applied to estimate the global Q-function to mitigate the impact of imperfect information sharing. We establish non-asymptotic global convergence guarantees under exact policy evaluation, where the rates are nearly independent of the size of the state-action space and illuminate the impacts of network size and connectivity. To the best of our knowledge, this is the first time that near dimension-free global convergence is established for federated multi-task RL using policy optimization. We further go beyond the tabular setting by proposing a federated natural actor critic (NAC) method for multi-task RL with function approximation, and establish its finite-time sample complexity taking the errors of function approximation into account.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- The Sample-Communication Complexity Trade-off in Federated Q-LearningSudeep Salgia, Yuejie ChiNeurIPS 2024 · 被引用 10 次
- Regret-Optimal Q-Learning with Low Cost for Single-Agent and Federated Reinforcement LearningHaochen Zhang, Zhong Zheng, Lingzhou XueNeurIPS 2025 · 被引用 3 次
- Constrained Multi-Objective Reinforcement Learning with Max-Min CriterionGiseung Park, Hyunyoung Nam, Woohyeon Byeon, Amir Leshem 等ICML 2026
- Exploration from a Primal-Dual Lens: Value-Incentivized Actor-Critic Methods for Sample-Efficient Online RLTong Yang, Bo Dai, Lin Xiao, Yuejie ChiNeurIPS 2025
它引用的顶会 Paper14
- Optimizing Federated Learning on Non-IID Data with Reinforcement LearningHao Wang, Zakhary Kaplan, Di Niu, Baochun LiINFOCOM 2020 · 被引用 1,002 次
- On the Global Convergence Rates of Softmax Policy Gradient MethodsJincheng Mei, Chenjun Xiao, Csaba Szepesvári, Dale SchuurmansICML 2020 · 被引用 349 次
- Neural Policy Gradient Methods: Global Optimality and Rates of ConvergenceLingxiao Wang, Qi Cai, Zhuoran Yang, Zhaoran WangICLR 2020 · 被引用 270 次
- Maximum Entropy RL (Provably) Solves Some Robust RL ProblemsBenjamin Eysenbach, Sergey LevineICLR 2022 · 被引用 244 次
- Adaptive Trust Region Policy Optimization: Global Convergence and Faster Rates for Regularized MDPsLior Shani, Yonathan Efroni, Shie MannorAAAI 2020 · 被引用 201 次
相关 Paper
- Improved Communication Efficiency in Federated Natural Policy Gradient via ADMM-based Gradient UpdatesGuangchen Lan, Han Wang, James Anderson, Christopher G. Brinton 等NeurIPS 2023 · 被引用 32 次
- A Sharper Global Convergence Analysis for Average Reward Reinforcement Learning via an Actor-Critic ApproachSwetha Ganesh, Washim Uddin Mondal, Vaneet AggarwalICML 2025
- Decentralized TD Tracking with Linear Function Approximation and its Finite-Time AnalysisGang Wang, Songtao Lu, Georgios B. Giannakis, Gerald Tesauro 等NeurIPS 2020 · 被引用 30 次
- Convergence of Policy Gradient for Entropy Regularized MDPs with Neural Network Approximation in the Mean-Field RegimeJames-Michael Leahy, Bekzhan Kerimkulov, David Siska, Lukasz SzpruchICML 2022 · 被引用 23 次
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar 等AAAI 2021 · 被引用 93 次
