Are Two Heads the Same as One? Identifying Disparate Treatment in Fair Neural Networks
Michael Lohaus, Matthäus Kleindessner, Krishnaram Kenthapadi, Francesco Locatello, Chris Russell
Abstract
We show that deep networks trained to satisfy demographic parity often do so through a form of race or gender awareness, and that the more we force a network to be fair, the more accurately we can recover race or gender from the internal state of the network. Based on this observation, we investigate an alternative fairness approach: we add a second classification head to the network to explicitly predict the protected attribute (such as race or gender) alongside the original task. After training the two-headed network, we enforce demographic parity by merging the two heads, creating a network with the same architecture as the original network. We establish a close relationship between existing approaches and our approach by showing (1) that the decisions of a fair classifier are well-approximated by our approach, and (2) that an unfair and optimally accurate classifier can be recovered from a fair classifier and our second head predicting the protected attribute. We use our explicit formulation to argue that the existing fairness approaches, just as ours, demonstrate disparate treatment and that they are likely to be unlawful in a wide range of scenarios under US law.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext a046f067-c9bf-4b2f-a75c-1fa9f87bdd5aCited by top-tier papers1
Ask how each one uses itBuilds on10
- Searching for MobileNetV3Andrew Howard, Ruoming Pang, Hartwig Adam, Quoc V. Le et al.ICCV 2019 · 9,163 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Racial Faces in the Wild: Reducing Racial Bias by Information Maximization Adaptation NetworkMei Wang, Weihong Deng, Jiani Hu, Xunqiang Tao et al.ICCV 2019 · 379 citations
- Overlearning Reveals Sensitive AttributesCongzheng Song, Vitaly ShmatikovICLR 2020 · 177 citations
- Too Relaxed to Be FairMichael Lohaus, Michaël Perrot, Ulrike von LuxburgICML 2020 · 80 citations
Related papers
- Learning Disentangled Representation for Fair Facial Attribute Classification via Fairness-aware Information AlignmentSungho Park, Sunhee Hwang, Dohyung Kim, Hyeran ByunAAAI 2021 · 68 citations
- The Disparate Benefits of Deep EnsemblesKajetan Schweighofer, Adrián Arnaiz-Rodríguez, Sepp Hochreiter, Nuria OliverICML 2025
- Fair Classification by Direct Intervention on Operating CharacteristicsKevin Jiang, Edgar DobribanICLR 2026
- Causal Context Connects Counterfactual Fairness to Robust Prediction and Group FairnessJacy Reese Anthis, Victor VeitchNeurIPS 2023 · 26 citations
- Fairness for Image Generation with Uncertain Sensitive AttributesAjil Jalal, Sushrut Karmalkar, Jessica Hoffmann, Alex Dimakis et al.ICML 2021 · 42 citations
