Characterizing Fairness Over the Set of Good Models Under Selective Labels
Amanda Coston, Ashesh Rambachan, Alexandra Chouldechova
Abstract
Algorithmic risk assessments are used to inform decisions in a wide variety of high-stakes settings. Often multiple predictive models deliver similar overall performance but differ markedly in their predictions for individual cases, an empirical phenomenon known as the "Rashomon Effect." These models may have different properties over various groups, and therefore have different predictive fairness properties. We develop a framework for characterizing predictive fairness properties over the set of models that deliver similar overall performance, or "the set of good models." Our framework addresses the empirically relevant challenge of selectively labelled data in the setting where the selection decision and outcome are unconfounded given the observed data features. Our framework can be used to 1) replace an existing model with one that has better fairness properties; or 2) audit for predictive bias. We illustrate these uses cases on a real-world credit-scoring task and a recidivism prediction task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers25
- Rashomon Capacity: A Metric for Predictive Multiplicity in ClassificationHsiang Hsu, Flávio P. CalmonNeurIPS 2022 · 65 citations
- Predictive Multiplicity in Probabilistic ClassificationJamelle Watson-Daniels, David C. Parkes, Berk UstunAAAI 2023 · 58 citations
- A Path to Simpler Models Starts With NoiseLesia Semenova, Harry Chen, Ronald Parr, Cynthia RudinNeurIPS 2023 · 41 citations
- What's the Harm? Sharp Bounds on the Fraction Negatively Affected by TreatmentNathan KallusNeurIPS 2022 · 40 citations
- Characterizing the risk of fairwashingUlrich Aïvodji, Hiromi Arai, Sébastien Gambs, Satoshi HaraNeurIPS 2021 · 35 citations
Builds on1
Related papers
- RashomonGB: Analyzing the Rashomon Effect and Mitigating Predictive Multiplicity in Gradient BoostingHsiang Hsu, Ivan Brugere, Shubham Sharma, Freddy Lécué et al.NeurIPS 2024 · 13 citations
- Exploring the cloud of feature interaction scores in a Rashomon setSichao Li, Rong Wang, Quanling Deng, Amanda S. BarnardICLR 2024 · 9 citations
- The Double-Edged Nature of the Rashomon Set for Trustworthy Machine LearningEthan Hsu, Harry Chen, Chudi Zhong, Lesia SemenovaICML 2026 · 1 citation
- Individual Arbitrariness and Group FairnessCarol Xuan Long, Hsiang Hsu, Wael Alghamdi, Flávio P. CalmonNeurIPS 2023 · 16 citations
- People Perceive Algorithmic Assessments as Less Fair and Trustworthy Than Identical Human AssessmentsLillio Mok, Sasha Nanda, Ashton AndersonCSCW 2023 · 18 citations
