Learning Fair Policies in Decentralized Cooperative Multi-Agent Reinforcement Learning
Matthieu Zimmer, Claire Glanois, Umer Siddique, Paul Weng
摘要
We consider the problem of learning fair policies in (deep) cooperative multi-agent reinforcement learning (MARL). We formalize it in a principled way as the problem of optimizing a welfare function that explicitly encodes two important aspects of fairness: efficiency and equity. We provide a theoretical analysis of the convergence of policy gradient for this problem. As a solution method, we propose a novel neural network architecture, which is composed of two sub-networks specifically designed for taking into account these two aspects of fairness. In experiments, we demonstrate the importance of the two sub-networks for fair optimization. Our overall approach is general as it can accommodate any (sub)differentiable welfare function. Therefore, it is compatible with various notions of fairness that have been proposed in the literature (e.g., lexicographic maximin, generalized Gini social welfare function, proportional fairness). Our method is generic and can be implemented in various MARL settings: centralized training and decentralized execution, or fully decentralized. Finally, we experimentally validate our approach in various domains and show that it can perform much better than previous methods, both in terms of efficiency and equity.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper14
- Two-sided fairness in rankings via Lorenz dominanceVirginie Do, Sam Corbett-Davies, Jamal Atif, Nicolas UsunierNeurIPS 2021 · 被引用 64 次
- Optimizing Generalized Gini Indices for Fairness in RankingsVirginie Do, Nicolas UsunierSIGIR 2022 · 被引用 19 次
- Cooperative Multi-Agent Fairness and Equivariant PoliciesNiko A. Grupen, Bart Selman, Daniel D. LeeAAAI 2022 · 被引用 15 次
- Dealing with Non-Stationarity in MARL via Trust-Region DecompositionWenhao Li, Xiangfeng Wang, Bo Jin, Junjie Sheng 等ICLR 2022 · 被引用 14 次
- Fair Off-Policy Learning from Observational DataDennis Frauen, Valentyn Melnychuk, Stefan FeuerriegelICML 2024 · 被引用 11 次
它引用的顶会 Paper2
- Weighted QMIX: Expanding Monotonic Value Function Factorisation for Deep Multi-Agent Reinforcement LearningTabish Rashid, Gregory Farquhar, Bei Peng, Shimon WhitesonNeurIPS 2020 · 被引用 1,960 次
- Learning Fair Policies in Multi-Objective (Deep) Reinforcement Learning with Average and Discounted RewardsUmer Siddique, Paul Weng, Matthieu ZimmerICML 2020 · 被引用 1 次
相关 Paper
- DIFFER: Decomposing Individual Reward for Fair Experience Replay in Multi-Agent Reinforcement LearningXunhan Hu, Jian Zhao, Wengang Zhou, Ruili Feng 等NeurIPS 2023 · 被引用 5 次
- Fair Resource Allocation in Multi-Task LearningHao Ban, Kaiyi JiICML 2024 · 被引用 41 次
- Specification-Guided Learning of Nash Equilibria with High Social WelfareKishor Jothimurugan, Suguman Bansal, Osbert Bastani, Rajeev AlurCAV 2022 · 被引用 9 次
- Achieving Fairness in Multi-Agent MDP Using Reinforcement LearningPeizhong Ju, Arnob Ghosh, Ness B. ShroffICLR 2024 · 被引用 8 次
- Navigating the Social Welfare Frontier: Portfolios for Multi-objective Reinforcement LearningCheol Woo Kim, Jai Moondra, Shresth Verma, Madeleine Pollack 等ICML 2025
