Enhancing Cooperative Multi-Agent Reinforcement Learning with State Modelling and Adversarial Exploration
Andreas Kontogiannis, Konstantinos Papathanasiou, Yi Shen, Giorgos Stamou, Michael M. Zavlanos, George A. Vouros
Abstract
Learning to cooperate in distributed partially observable environments with no communication abilities poses significant challenges for multiagent deep reinforcement learning (MARL). This paper addresses key concerns in this domain, focusing on inferring state representations from individual agent observations and leveraging these representations to enhance agents' exploration and collaborative task execution policies. To this end, we propose a novel state modelling framework for cooperative MARL, where agents infer meaningful belief representations of the nonobservable state, with respect to optimizing their own policies, while filtering redundant and less informative joint state information. Building upon this framework, we propose the MARL SMPE 2 algorithm. In SMPE 2 , agents enhance their own policy's discriminative abilities under partial observability, explicitly by incorporating their beliefs into the policy network, and implicitly by adopting an adversarial type of exploration policies which encourages agents to discover novel, highvalue states while improving the discriminative abilities of others. Experimentally, we show that SMPE 2 outperforms state-of-the-art MARL algorithms in complex fully cooperative tasks from the MPE, LBF, and RWARE benchmarks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 790519a4-3d2b-41c7-a57f-4527bef4ef69Cited by top-tier papers2
- Diffusing to Coordinate: Efficient Online Multi-Agent Diffusion PoliciesZhuoran Li, Hai Zhong, Xun Wang, Qingxin Xia et al.ICML 2026 · 2 citations
- Sparse Topology-Aware Pairwise Scoring for Large-Scale Multi-Agent Reinforcement LearningZhibo Deng, Feng Liang, Yong Zhang, Xiaoxi Zhang et al.ICML 2026
Builds on18
- Multi-Agent Reinforcement Learning is a Sequence Modeling ProblemMuning Wen, Jakub Grudzien Kuba, Runji Lin, Weinan Zhang et al.NeurIPS 2022 · 408 citations
- Shared Experience Actor-Critic for Multi-Agent Reinforcement LearningFilippos Christianos, Lukas Schäfer, Stefano V. AlbrechtNeurIPS 2020 · 238 citations
- Celebrating Diversity in Shared Multi-Agent Reinforcement LearningChenghao Li, Tonghan Wang, Chengjie Wu, Qianchuan Zhao et al.NeurIPS 2021 · 224 citations
- Scaling Multi-Agent Reinforcement Learning with Selective Parameter SharingFilippos Christianos, Georgios Papoudakis, Arrasy Rahman, Stefano V. AlbrechtICML 2021 · 165 citations
- Value-Decomposition Multi-Agent Actor-CriticsJianyu Su, Stephen C. Adams, Peter A. BelingAAAI 2021 · 140 citations
Related papers
- MA2E: Addressing Partial Observability in Multi-Agent Reinforcement Learning with Masked Auto-EncoderSehyeok Kang, Yongsik Lee, Gahee Kim, Song Chong et al.ICLR 2025
- Think How Your Teammates Think: Active Inference Can Benefit Decentralized ExecutionHao Wu, Shoucheng Song, Chang Yao, Sheng Han et al.AAAI 2026 · 1 citation
- GRDC: A Unified Graph-Driven Framework for Role Discovery and Communication in Multi-Agent Reinforcement LearningZihong Gao, Hongjian Liang, Yuanhui Hao, Lei Hao et al.AAAI 2026
- Variational Recurrent Models for Solving Partially Observable Control TasksDongqi Han, Kenji Doya, Jun TaniICLR 2020 · 75 citations
- LLM-Guided Communication for Cooperative Multi-Agent Reinforcement LearningSangjun Bae, Yisak Park, Sanghyeon Lee, Seungyul HanICML 2026 · 2 citations
