Factored Policy Gradients: Leveraging Structure for Efficient Learning in MOMDPs
Thomas Spooner, Nelson Vadori, Sumitra Ganesh
Abstract
Policy gradient methods can solve complex tasks but often fail when the dimensionality of the action-space or objective multiplicity grow very large. This occurs, in part, because the variance on score-based gradient estimators scales quadratically. In this paper, we address this problem through a factor baseline which exploits independence structure encoded in a novel action-target influence network. Factored policy gradients (FPGs), which follow, provide a common framework for analysing key state-of-the-art algorithms, are shown to generalise traditional policy gradients, and yield a principled way of incorporating prior knowledge of a problem domain's generative processes. We provide an analysis of the proposed estimator and identify the conditions under which variance is reduced. The algorithmic aspects of FPGs are discussed, including optimal policy factorisation, as characterised by minimum biclique coverings, and the implications for the bias-variance trade-off of incorrectly specifying the network. Finally, we demonstrate the performance advantages of our algorithm on large-scale bandit and traffic intersection problems, providing a novel contribution to the latter in the form of a spatial approximation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Leveraging Factored Action Spaces for Efficient Offline Reinforcement Learning in HealthcareShengpu Tang, Maggie Makar, Michael W. Sjoding, Finale Doshi-Velez et al.NeurIPS 2022 · 63 citations
- Disentangled Unsupervised Skill Discovery for Efficient Hierarchical Reinforcement LearningJiaheng Hu, Zizhao Wang, Peter Stone, Roberto Martín-MartínNeurIPS 2024 · 21 citations
- Distributed Influence-Augmented Local Simulators for Parallel MARL in Large Networked SystemsMiguel Suau, Jinke He, Mustafa Mert Çelikok, Matthijs T. J. Spaan et al.NeurIPS 2022 · 2 citations
Builds on2
Related papers
- Settling the Variance of Multi-Agent Policy GradientsJakub Grudzien Kuba, Muning Wen, Linghui Meng, Shangding Gu et al.NeurIPS 2021 · 121 citations
- PAGE-PG: A Simple and Loopless Variance-Reduced Policy Gradient Method with Probabilistic Gradient EstimationMatilde Gargiani, Andrea Zanelli, Andrea Martinelli, Tyler H. Summers et al.ICML 2022 · 17 citations
- FACMAC: Factored Multi-Agent Centralised Policy GradientsBei Peng, Tabish Rashid, Christian Schröder de Witt, Pierre-Alexandre Kamienny et al.NeurIPS 2021 · 399 citations
- Optimal Estimation of Policy Gradient via Double Fitted IterationChengzhuo Ni, Ruiqi Zhang, Xiang Ji, Xuezhou Zhang et al.ICML 2022 · 1 citation
- On the Convergence and Sample Efficiency of Variance-Reduced Policy Gradient MethodJunyu Zhang, Chengzhuo Ni, Zheng Yu, Csaba Szepesvári et al.NeurIPS 2021 · 87 citations
