Feature Importance Disparities for Data Bias Investigations
Peter W. Chang, Leor Fishman, Seth Neel
摘要
It is widely held that one cause of downstream bias in classifiers is bias present in the training data. Rectifying such biases may involve context-dependent interventions such as training separate models on subgroups, removing features with bias in the collection process, or even conducting real-world experiments to ascertain sources of bias. Despite the need for such data bias investigations, few automated methods exist to assist practitioners in these efforts. In this paper, we present one such method that given a dataset consisting of protected and unprotected features, outcomes , and a regressor that predicts given , outputs a tuple , with the following property: corresponds to a subset of the training dataset , such that the feature has much larger (or smaller) influence in the subgroup , than on the dataset overall, which we call feature importance disparity (FID). We show across datasets and common feature importance methods of broad interest to the machine learning community that we can efficiently find subgroups with large FID values even over exponentially large subgroup classes and in practice these groups correspond to subgroups with potentially serious bias issues as measured by standard fairness metrics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper2
相关 Paper
- Understanding challenges to the interpretation of disaggregated evaluations of algorithmic fairnessStephen Pfohl, Natalie Harris, Chirag Nagpal, David Madras 等NeurIPS 2025 · 被引用 9 次
- Causal Explanations for Disparate Trends: Where and Why?Tal Blau, Brit Youngmann, Anna Fariha, Yuval MoskovitchSIGMOD 2026 · 被引用 2 次
- Mitigating Subgroup Unfairness in Machine Learning Classifiers: A Data-Driven ApproachYin Lin, Samika Gupta, H. V. JagadishICDE 2024 · 被引用 5 次
- Fair Infinitesimal Jackknife: Mitigating the Influence of Biased Training Data Points Without RefittingPrasanna Sattigeri, Soumya Ghosh, Inkit Padhi, Pierre L. Dognin 等NeurIPS 2022 · 被引用 36 次
- FIPSER: Improving Fairness Testing of DNN by Seed PrioritizationJunwei Chen, Yueling Zhang, Lingfeng Zhang, Min Zhang 等ASE 2024 · 被引用 2 次
