Fairness Aware Reinforcement Learning via Proximal Policy Optimization
Gabriele La Malfa, Jie M. Zhang, Michael Luck, Elizabeth Black
Abstract
Fairness in multi-agent systems (MAS) focuses on equitable reward distribution among agents in scenarios involving sensitive attributes such as race, gender, or socioeconomic status. This paper introduces fairness in Proximal Policy Optimization (PPO) with a penalty term derived from a fairness definition such as demographic parity, counterfactual fairness, or conditional statistical parity. The proposed method, which we call Fair-PPO, balances reward maximisation with fairness by integrating two penalty components: a retrospective component that minimises disparities in past outcomes and a prospective component that ensures fairness in future decision-making. We evaluate our approach in two games: the Allelopathic Harvest, a cooperative and competitive MAS focused on resource collection, where some agents possess a sensitive attribute, and HospitalSim, a hospital simulation, in which agents coordinate the operations of hospital patients with different mobility and priority needs. Experiments show that Fair-PPO achieves fairer policies than PPO across the fairness metrics and, through the retrospective and prospective penalty components, reveals a wide spectrum of strategies to improve fairness; at the same time, its performance pairs with that of state-of-the-art fair reinforcement-learning algorithms. Fairness comes at the cost of reduced efficiency, but does not compromise equality among the overall population (Gini index). These findings underscore the potential of Fair-PPO to address fairness challenges in MAS. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 4d8a867d-de2b-49c4-80af-9b68b912a879Builds on6
- Achieving Fairness in the Stochastic Multi-Armed Bandit ProblemVishakha Patil, Ganesh Ghalme, Vineet Nair, Y. NarahariAAAI 2020 · 131 citations
- Learning Fair Policies in Decentralized Cooperative Multi-Agent Reinforcement LearningMatthieu Zimmer, Claire Glanois, Umer Siddique, Paul WengICML 2021 · 76 citations
- Bringing Fairness to Actor-Critic Reinforcement Learning for Network Utility OptimizationJingdi Chen, Yimeng Wang, Tian LanINFOCOM 2021 · 23 citations
- An Efficient Algorithm for Fair Multi-Agent Multi-Armed Bandit with Low RegretMatthew Jones, Huy L. Nguyen, Thy Dinh NguyenAAAI 2023 · 11 citations
- Achieving Fairness in Multi-Agent MDP Using Reinforcement LearningPeizhong Ju, Arnob Ghosh, Ness B. ShroffICLR 2024 · 8 citations
Related papers
- Learning Diverse Risk Preferences in Population-Based Self-PlayYuhua Jiang, Qihan Liu, Xiaoteng Ma, Chenghao Li et al.AAAI 2024 · 8 citations
- Coordinated Proximal Policy OptimizationZifan Wu, Chao Yu, Deheng Ye, Junge Zhang et al.NeurIPS 2021 · 73 citations
- Group Meritocratic Fairness in Linear Contextual BanditsRiccardo Grazzi, Arya Akhavan, John Isak Texas Falk, Leonardo Cella et al.NeurIPS 2022 · 12 citations
- FairICP: Encouraging Equalized Odds via Inverse Conditional PermutationYuheng Lai, Leying GuanICML 2025
- Projection-Based Constrained Policy OptimizationTsung-Yen Yang, Justinian Rosca, Karthik Narasimhan, Peter J. RamadgeICLR 2020 · 306 citations
