Lune

ICML2024顶会

Feature Importance Disparities for Data Bias Investigations

Peter W. Chang, Leor Fishman, Seth Neel

2024年份
3被引次数

摘要

It is widely held that one cause of downstream bias in classifiers is bias present in the training data. Rectifying such biases may involve context-dependent interventions such as training separate models on subgroups, removing features with bias in the collection process, or even conducting real-world experiments to ascertain sources of bias. Despite the need for such data bias investigations, few automated methods exist to assist practitioners in these efforts. In this paper, we present one such method that given a dataset XX consisting of protected and unprotected features, outcomes yy, and a regressor hh that predicts yy given XX, outputs a tuple (fj,g)(f_j, g), with the following property: gg corresponds to a subset of the training dataset (X,y)(X, y), such that the jthj^{th} feature fjf_j has much larger (or smaller) influence in the subgroup gg, than on the dataset overall, which we call feature importance disparity (FID). We show across 44 datasets and 44 common feature importance methods of broad interest to the machine learning community that we can efficiently find subgroups with large FID values even over exponentially large subgroup classes and in practice these groups correspond to subgroups with potentially serious bias issues as measured by standard fairness metrics.

问问这篇 Paper

智能体会读完全文。

Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。

可以从这些问题问起

智能体调用

Luneget_paper_fulltext

在 Lune 里问

免费开始,无需绑卡

它引用的顶会 Paper2

相关 Paper

黄昏的海面,两侧是细线勾勒的悬崖