Learning Fair Policies in Decentralized Cooperative Multi-Agent Reinforcement Learning
Matthieu Zimmer, Claire Glanois, Umer Siddique, Paul Weng
Abstract
We consider the problem of learning fair policies in (deep) cooperative multi-agent reinforcement learning (MARL). We formalize it in a principled way as the problem of optimizing a welfare function that explicitly encodes two important aspects of fairness: efficiency and equity. We provide a theoretical analysis of the convergence of policy gradient for this problem. As a solution method, we propose a novel neural network architecture, which is composed of two sub-networks specifically designed for taking into account these two aspects of fairness. In experiments, we demonstrate the importance of the two sub-networks for fair optimization. Our overall approach is general as it can accommodate any (sub)differentiable welfare function. Therefore, it is compatible with various notions of fairness that have been proposed in the literature (e.g., lexicographic maximin, generalized Gini social welfare function, proportional fairness). Our method is generic and can be implemented in various MARL settings: centralized training and decentralized execution, or fully decentralized. Finally, we experimentally validate our approach in various domains and show that it can perform much better than previous methods, both in terms of efficiency and equity.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 65f8902e-566f-4bd8-a544-be909e125c32Cited by top-tier papers14
- Two-sided fairness in rankings via Lorenz dominanceVirginie Do, Sam Corbett-Davies, Jamal Atif, Nicolas UsunierNeurIPS 2021 · 64 citations
- Optimizing Generalized Gini Indices for Fairness in RankingsVirginie Do, Nicolas UsunierSIGIR 2022 · 19 citations
- Cooperative Multi-Agent Fairness and Equivariant PoliciesNiko A. Grupen, Bart Selman, Daniel D. LeeAAAI 2022 · 15 citations
- Dealing with Non-Stationarity in MARL via Trust-Region DecompositionWenhao Li, Xiangfeng Wang, Bo Jin, Junjie Sheng et al.ICLR 2022 · 14 citations
- Fair Off-Policy Learning from Observational DataDennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICML 2024 · 11 citations
Builds on2
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 1,960 citations
- Learning Fair Policies in Multi-Objective (Deep) Reinforcement Learning with Average and Discounted RewardsUmer Siddique, Paul Weng, Matthieu ZimmerICML 2020 · 1 citation
Related papers
- DIFFER: Decomposing Individual Reward for Fair Experience Replay in Multi-Agent Reinforcement LearningXunhan Hu, Jian Zhao, Wengang Zhou, Ruili Feng et al.NeurIPS 2023 · 5 citations
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 41 citations
- Specification-Guided Learning of Nash Equilibria with High Social WelfareKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurCAV 2022 · 9 citations
- Achieving Fairness in Multi-Agent MDP Using Reinforcement LearningPeizhong Ju, Arnob Ghosh, Ness B. ShroffICLR 2024 · 8 citations
- Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement LearningCheol Woo Kim, Jai Moondra, Shresth Verma, Madeleine Pollack et al.ICML 2025
