Fairness Aware Reinforcement Learning via Proximal Policy Optimization
Gabriele La Malfa, Jie M. Zhang, Michael Luck, Elizabeth Black
摘要
Fairness in multi-agent systems (MAS) focuses on equitable reward distribution among agents in scenarios involving sensitive attributes such as race, gender, or socioeconomic status. This paper introduces fairness in Proximal Policy Optimization (PPO) with a penalty term derived from a fairness definition such as demographic parity, counterfactual fairness, or conditional statistical parity. The proposed method, which we call Fair-PPO, balances reward maximisation with fairness by integrating two penalty components: a retrospective component that minimises disparities in past outcomes and a prospective component that ensures fairness in future decision-making. We evaluate our approach in two games: the Allelopathic Harvest, a cooperative and competitive MAS focused on resource collection, where some agents possess a sensitive attribute, and HospitalSim, a hospital simulation, in which agents coordinate the operations of hospital patients with different mobility and priority needs. Experiments show that Fair-PPO achieves fairer policies than PPO across the fairness metrics and, through the retrospective and prospective penalty components, reveals a wide spectrum of strategies to improve fairness; at the same time, its performance pairs with that of state-of-the-art fair reinforcement-learning algorithms. Fairness comes at the cost of reduced efficiency, but does not compromise equality among the overall population (Gini index). These findings underscore the potential of Fair-PPO to address fairness challenges in MAS. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- Achieving Fairness in the Stochastic Multi-Armed Bandit ProblemVishakha Patil, Ganesh Ghalme, Vineet Nair, Y. NarahariAAAI 2020 · 被引用 131 次
- Learning Fair Policies in Decentralized Cooperative Multi-Agent Reinforcement LearningMatthieu Zimmer, Claire Glanois, Umer Siddique, Paul WengICML 2021 · 被引用 76 次
- Bringing Fairness to Actor-Critic Reinforcement Learning for Network Utility OptimizationJingdi Chen, Yimeng Wang, Tian LanINFOCOM 2021 · 被引用 23 次
- An Efficient Algorithm for Fair Multi-Agent Multi-Armed Bandit with Low RegretMatthew Jones, Huy L. Nguyen, Thy Dinh NguyenAAAI 2023 · 被引用 11 次
- Achieving Fairness in Multi-Agent MDP Using Reinforcement LearningPeizhong Ju, Arnob Ghosh, Ness B. ShroffICLR 2024 · 被引用 8 次
相关 Paper
- Learning Diverse Risk Preferences in Population-Based Self-PlayYuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li 等AAAI 2024 · 被引用 8 次
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang 等NeurIPS 2021 · 被引用 73 次
- Group Meritocratic Fairness in Linear Contextual BanditsRiccardo Grazzi, Arya Akhavan, John Isak Texas Falk, Leonardo Cella 等NeurIPS 2022 · 被引用 12 次
- FairICP: Encouraging Equalized Odds via Inverse Conditional PermutationYuheng Lai, Leying GuanICML 2025
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 被引用 306 次
