In-Context Fully Decentralized Cooperative Multi-Agent Reinforcement Learning
Chao Li, Bingkun Bao, Yang Gao
Abstract
In this paper, we consider fully decentralized cooperative multi-agent reinforcement learning, where each agent has access only to the states, its local actions, and the shared rewards. The absence of information about other agents’ actions typically leads to the non-stationarity problem during per-agent value function updates, and the relative overgeneralization issue during value function estimation. However, existing works fail to address both issues simultaneously, as they lack the capability to model the agents’ joint policy in a fully decentralized setting. To overcome this limitation, we propose a simple yet effective method named Return-Aware Context (RAC). RAC formalizes the dynamically changing task, as locally perceived by each agent, as a contextual Markov Decision Process (MDP), and addresses both non-stationarity and relative overgeneralization through return-aware context modeling. Specifically, the contextual MDP attributes the non-stationary local dynamics of each agent to switches between contexts, each corresponding to a distinct joint policy. Then, based on the assumption that the joint policy changes only between episodes, RAC distinguishes different joint policies by the training episodic return and constructs contexts using discretized episodic return values. Accordingly, RAC learns a context-based value function for each agent to address the non-stationarity issue during value function updates. For value function estimation, an individual optimistic marginal value is constructed to encourage the selection of optimal joint actions, thereby mitigating the relative overgeneralization problem. Experimentally, we evaluate RAC on various cooperative tasks (including matrix game, predator and prey, and SMAC), and its significant performance validates its effectiveness.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 60eab0ab-e6e6-4b41-9601-933540c199b5Builds on4
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- UneVEn: Universal Value Exploration for Multi-Agent Reinforcement LearningTarun Gupta, Anuj Mahajan, Bei Peng, Wendelin Boehmer et al.ICML 2021 · 59 citations
- I2Q: A Fully Decentralized Q-Learning AlgorithmJiechuan Jiang, Zongqing LuNeurIPS 2022 · 34 citations
- GAT-MF: Graph Attention Mean Field for Very Large Scale Multi-Agent Reinforcement LearningQianyue Hao, Wenzhen Huang, Tao Feng, Jian Yuan et al.KDD 2023 · 18 citations
Related papers
- Multi-agent In-context Coordination via Decentralized Memory RetrievalTao Jiang, Zichuan Lin, Lihe Li, Yi-Chen Li et al.AAAI 2026 · 1 citation
- Decentralized Q-learning in Zero-sum Markov GamesMuhammed O. Sayin, Kaiqing Zhang, David S. Leslie, Tamer Basar et al.NeurIPS 2021 · 105 citations
- Potentially Optimal Joint Actions Recognition for Cooperative Multi-Agent Reinforcement LearningChang Huang, Shatong Zhu, Junqiao Zhao, Hongtu Zhou et al.ICLR 2026
- More Centralized Training, Still Decentralized Execution: Multi-Agent Conditional Policy FactorizationJiangxing Wang, Deheng Ye, Zongqing LuICLR 2023 · 5 citations
- Multi-Agent Reinforcement Learning with General Utilities via Decentralized Shadow Reward Actor-CriticJunyu Zhang, Amrit Singh Bedi, Mengdi Wang, Alec KoppelAAAI 2022 · 7 citations
