Measuring Fairness of Text Classifiers via Prediction Sensitivity
Satyapriya Krishna, Rahul Gupta, Apurv Verma, Jwala Dhamala, Yada Pruksachatkun, Kai-Wei Chang
Abstract
With the rapid growth in language processing applications, fairness has emerged as an important consideration in data-driven solutions. Although various fairness definitions have been explored in the recent literature, there is lack of consensus on which metrics most accurately reflect the fairness of a system. In this work, we propose a new formulation – accumulated prediction sensitivity, which measures fairness in machine learning models based on the model’s prediction sensitivity to perturbations in input features. The metric attempts to quantify the extent to which a single prediction depends on a protected attribute, where the protected attribute encodes the membership status of an individual in a protected group. We show that the metric can be theoretically linked with a specific notion of group fairness (statistical parity) and individual fairness. It also correlates well with humans’ perception of fairness. We conduct experiments on two text classification datasets – Jigsaw Toxicity, and Bias in Bios, and evaluate the correlations between metrics and manual annotations on whether the model produced a fair outcome. We observe that the proposed fairness metric based on prediction sensitivity is statistically significantly more correlated with human annotation than the existing counterfactual fairness metric.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 8d07bd3f-ecef-4111-8213-323263e56943Cited by top-tier papers3
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett et al.EMNLP 2024 · 4 citations
- Certifying Counterfactual Bias in LLMsIsha Chaudhary, Qian Hu, Manoj Kumar, Morteza Ziyadi et al.ICLR 2025 · 3 citations
- Self-Regularization with Sparse Autoencoders for Controllable LLM-based ClassificationXuansheng Wu, Wenhao Yu, Xiaoming Zhai, Ninghao LiuKDD 2025
Builds on7
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Explaining Black Box Predictions and Unveiling Data Artifacts through Influence FunctionsXiaochuang Han, Byron C. Wallace, Yulia TsvetkovACL 2020 · 91 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston et al.EMNLP 2020 · 7 citations
Related papers
- Fair Without Leveling Down: A New Intersectional Fairness DefinitionGaurav Maheshwari, Aurélien Bellet, Pascal Denis, Mikaela KellerEMNLP 2023 · 3 citations
- Bayes-Optimal Fair Classification with Multiple Sensitive FeaturesYi Yang, Yinghui Huang, Xiangyu ChangAAAI 2026 · 2 citations
- Rényi Fair InferenceSina Baharlouei, Maher Nouiehed, Ahmad Beirami, Meisam RazaviyaynICLR 2020 · 69 citations
- Two Simple Ways to Learn Individual Fairness Metrics from DataDebarghya Mukherjee, Mikhail Yurochkin, Moulinath Banerjee, Yuekai SunICML 2020 · 109 citations
- Pairwise Fairness for Ranking and RegressionHarikrishna Narasimhan, Andrew Cotter, Maya R. Gupta, Serena Lutong WangAAAI 2020 · 125 citations
