Decentralized TD Tracking with Linear Function Approximation and its Finite-Time Analysis
Gang Wang, Songtao Lu, Georgios B. Giannakis, Gerald Tesauro, Jian Sun
Abstract
The present contribution deals with decentralized policy evaluation in multi-agent Markov decision processes using temporal-difference (TD) methods with linear function approximation for scalability. The agents cooperate to estimate the value function of such a process by observing continual state transitions of a shared environment over the graph of interconnected nodes (agents), along with locally private rewards. Different from existing consensus-type TD algorithms, the approach here develops a simple decentralized TD tracker by wedding TD learning with gradient tracking techniques. The non-asymptotic properties of the novel TD tracker are established for both independent and identically distributed (i.i.d.) as well as Markovian transitions through a unifying multistep Lyapunov analysis. In contrast to the prior art, the novel algorithm forgoes the limiting error bounds on the number of agents, which endows it with performance comparable to that of centralized TD methods that are the sharpest known to date.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0c9e4527-a82e-4a76-8246-9b90ab19e1ccCited by top-tier papers8
- Federated Reinforcement Learning: Linear Speedup Under Markovian SamplingSajad Khodadadian, Pranay Sharma, Gauri Joshi, Siva Theja MaguluriICML 2022 · 46 citations
- Taming Communication and Sample Complexities in Decentralized Policy Evaluation for Cooperative Multi-Agent Reinforcement LearningXin Zhang, Zhuqing Liu, Jia Liu, Zhengyuan Zhu et al.NeurIPS 2021 · 36 citations
- Sample and Communication-Efficient Decentralized Actor-Critic Algorithms with Finite-Time AnalysisZiyi Chen, Yi Zhou, Rong-Rong Chen, Shaofeng ZouICML 2022 · 35 citations
- Federated Q-Learning: Linear Regret Speedup with Low Communication CostZhong Zheng, Fengyu Gao, Lingzhou Xue, Jing YangICLR 2024 · 21 citations
- The Sample-Communication Complexity Trade-off in Federated Q-LearningSudeep Salgia, Yuejie ChiNeurIPS 2024 · 10 citations
Builds on2
Related papers
- Multi-Agent Reinforcement Learning in Stochastic Networked SystemsYiheng Lin, Guannan Qu, Longbo Huang, Adam WiermanNeurIPS 2021 · 55 citations
- MDPGT: Momentum-Based Decentralized Policy Gradient TrackingZhanhong Jiang, Xian Yeow Lee, Sin Yong Tan, Kai Liang Tan et al.AAAI 2022 · 11 citations
- Federated Natural Policy Gradient and Actor Critic Methods for Multi-task Reinforcement LearningTong Yang, Shicong Cen, Yuting Wei, Yuxin Chen et al.NeurIPS 2024 · 14 citations
- Non-Asymptotic Analysis for Two Time-scale TDC with General Smooth Function ApproximationYue Wang, Shaofeng Zou, Yi ZhouNeurIPS 2021 · 12 citations
- Decentralized Policy Gradient Descent Ascent for Safe Multi-Agent Reinforcement LearningSongtao Lu, Kaiqing Zhang, Tianyi Chen, Tamer Basar et al.AAAI 2021 · 93 citations
