Decomposing Temporal Equilibrium Strategy for Coordinated Distributed Multi-Agent Reinforcement Learning
Chenyang Zhu, Wen Si, Jinyu Zhu, Zhihao Jiang
Abstract
The increasing demands for system complexity and robustness have prompted the integration of temporal logic into Multi-Agent Reinforcement Learning (MARL) to address tasks with non-Markovian properties. However, incorporating non-Markovian properties introduces additional computational complexities, as agents are required to integrate historical data into their decision-making process. Also, optimizing strategies within a multi-agent environment presents significant challenges due to the exponential growth of the state space with the number of agents. In this study, we introduce an innovative hierarchical MARL framework that synthesizes temporal equilibrium strategies through parity games and subsequently encodes them as individual reward machines for MARL coordination. More specifically, we reduce the strategy synthesis problem into an emptiness problem concerning parity games with optimized states and transitions. Following this synthesis step, the temporal equilibrium strategy is decomposed into individual reward machines for decentralized MARL. Theoretical proofs are provided to verify the consistency of the Nash equilibrium between the parallel composition of decomposed strategies and the original strategy. Empirical evidence confirms the efficacy of the proposed synthesis technique, showcasing its ability to reduce state space compared to the state-of-the-art tool. Furthermore, our study highlights the superior performance of the distributed MARL paradigm over centralized approaches when deploying decomposed strategies.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9c94b5e7-8fab-41c6-9e1c-3720df7c06d1Builds on1
Related papers
- T4NMTD: Transition-Centric Reinforcement Learning for Non-Markovian Task DecompositionRuixuan Miao, Xu Lu, Cong Tian, Bin Yu et al.AAAI 2026
- HyPOLE: Hyperproperty-Guided Multi-Agent Reinforcement Learning under Partial ObservationArshia Rafieioskouei, Tzu-Han Hsu, Matthew Lucas, Borzoo BonakdarpourICML 2026
- Bi-Level Actor-Critic for Multi-Agent CoordinationHaifeng Zhang, Weizhe Chen, Zeren Huang, Minne Li et al.AAAI 2020 · 113 citations
- Leveraging Partial Symmetry for Multi-Agent Reinforcement LearningXin Yu, Rongye Shi, Pu Feng, Yongkai Tian et al.AAAI 2024 · 24 citations
- From Debate to Equilibrium: Belief‑Driven Multi‑Agent LLM Reasoning via Bayesian Nash EquilibriumXie Yi, Zhanke Zhou, Chentao Cao, Qiyu Niu et al.ICML 2025
