Asynchronous Actor-Critic for Multi-Agent Reinforcement Learning
Yuchen Xiao, Weihao Tan, Christopher Amato
Abstract
Synchronizing decisions across multiple agents in realistic settings is problematic since it requires agents to wait for other agents to terminate and communicate about termination reliably. Ideally, agents should learn and execute asynchronously instead. Such asynchronous methods also allow temporally extended actions that can take different amounts of time based on the situation and action executed. Unfortunately, current policy gradient methods are not applicable in asynchronous settings, as they assume that agents synchronously reason about action selection at every time step. To allow asynchronous learning and decision-making, we formulate a set of asynchronous multi-agent actor-critic methods that allow agents to directly optimize asynchronous policies in three standard training paradigms: decentralized learning, centralized learning, and centralized training for decentralized execution. Empirical results (in simulation and hardware) in a variety of realistic domains demonstrate the superiority of our approaches in large multi-agent problems and validate the effectiveness of our algorithms for learning high-quality and asynchronous solutions.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers7
- True Knowledge Comes from Practice: Aligning Large Language Models with Embodied Environments via Reinforcement LearningWeihao Tan, Wentao Zhang, Shanqi Liu, Longtao Zheng et al.ICLR 2024 · 33 citations
- M³HF: Multi-agent Reinforcement Learning from Multi-phase Human Feedback of Mixed QualityZiyan Wang, Zhicheng Zhang, Fei Fang, Yali DuICML 2025
- Cradle: Empowering Foundation Agents towards General Computer ControlWeihao Tan, Wentao Zhang, Xinrun Xu, Haochong Xia et al.ICML 2025
- Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial ExplorationAndreas Kontogiannis, Konstantinos Papathanasiou, Yi Shen, Giorgos Stamou et al.ICML 2025
- Agent-Centric Actor-Critic for Asynchronous Multi-Agent Reinforcement LearningWhiyoung Jung, Sunghoon Hong, Deunsol Yoon, Kanghoon Lee et al.ICML 2025
Builds on9
- Emergent Tool Use From Multi-Agent AutocurriculaBowen Baker, Ingmar Kanitscheider, Todor M. Markov, Yi Wu et al.ICLR 2020 · 751 citations
- DOP: Off-Policy Multi-Agent Decomposed Policy GradientsYihan Wang, Beining Han, Tonghan Wang, Heng Dong et al.ICLR 2021 · 208 citations
- Shapley Q-Value: A Local Reward Approach to Solve Global Reward GamesJianhong Wang, Yuan Zhang, Tae-Kyun Kim, Yunjie GuAAAI 2020 · 159 citations
- Learning Implicit Credit Assignment for Cooperative Multi-Agent Reinforcement LearningMeng Zhou, Ziyu Liu, Pengwei Sui, Yixuan Li et al.NeurIPS 2020 · 142 citations
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 140 citations
Related papers
- A Deeper Understanding of State-Based Critics in Multi-Agent Reinforcement LearningXueguang Lyu, Andrea Baisero, Yuchen Xiao, Christopher AmatoAAAI 2022 · 19 citations
- Communication-Efficient Actor-Critic Methods for Homogeneous Markov GamesDingyang Chen, Yile Li, Qi ZhangICLR 2022 · 11 citations
- Finite-Time Global Optimality Convergence in Deep Neural Actor-Critic Methods for Decentralized Multi-Agent Reinforcement LearningZhiyao Zhang, Myeung Suk Oh, Hairi, Ziyue Luo et al.ICML 2025
- Learning Decentralized LLM Collaboration with Multi-Agent Actor CriticShuo Liu, Tianle Chen, Ryan Amiri, Christopher AmatoICML 2026 · 6 citations
- Revisiting Some Common Practices in Cooperative Multi-Agent Reinforcement LearningWei Fu, Chao Yu, Zelai Xu, Jiaqi Yang et al.ICML 2022 · 49 citations
