Bias Detection via Maximum Subgroup Discrepancy
Jiri Nemecek, Mark Kozdoba, Illia Kryvoviaz, Tomás Pevný, Jakub Marecek
摘要
Bias evaluation is fundamental to trustworthy AI, both in terms of checking data quality and in terms of checking the outputs of AI systems.In testing data quality, for example, one may study the distance of a given dataset, viewed as a distribution, to a given ground-truth reference dataset.However, classical metrics, such as the Total Variation and the Wasserstein distances, are known to have high sample complexities and, therefore, may fail to provide a meaningful distinction in many practical scenarios.In this paper, we propose a new notion of distance, the Maximum Subgroup Discrepancy (MSD).In this metric, two distributions are close if, roughly, discrepancies are low for all feature subgroups.While the number of subgroups may be exponential, we show that the sample complexity is linear in the number of features, thus making it feasible for practical applications.Moreover, we provide a practical algorithm for evaluating the distance based on Mixedinteger optimization (MIO).We also note that the proposed distance is easily interpretable, thus providing clearer paths to fixing the biases once they have been identified.Finally, we describe a natural general bias detection framework, termed MSDD distances, and show that MSD aligns well with this framework.We empirically evaluate MSD by comparing it with other metrics and by demonstrating the above properties of MSD on real-world datasets. CCS Concepts Theory of computation Complexity theory and logic; Mathematics of computing Multivariate statistics.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper7
- Retiring Adult: New Datasets for Fair Machine LearningFrances Ding, Moritz Hardt, John Miller, Ludwig SchmidtNeurIPS 2021 · 被引用 671 次
- Learning from Failure: De-biasing Classifier from Biased ClassifierJun Hyun Nam, Hyuntak Cha, Sungsoo Ahn, Jaeho Lee 等NeurIPS 2020 · 被引用 428 次
- Learning De-biased Representations with Biased RepresentationsHyojin Bahng, Sanghyuk Chun, Sangdoo Yun, Jaegul Choo 等ICML 2020 · 被引用 332 次
- Fairness Improvement with Multiple Protected Attributes: How Far Are We?Zhenpeng Chen, Jie M. Zhang, Federica Sarro, Mark HarmanICSE 2024 · 被引用 33 次
- Amortized Projection Optimization for Sliced Wasserstein Generative ModelsKhai Nguyen, Nhat HoNeurIPS 2022 · 被引用 23 次
相关 Paper
- On the Maximal Local Disparity of Fairness-Aware ClassifiersJinqiu Jin, Haoxuan Li, Fuli FengICML 2024 · 被引用 5 次
- Discover and Mitigate Multiple Biased Subgroups in Image ClassifiersZeliang Zhang, Mingqian Feng, Zhiheng Li, Chenliang XuCVPR 2024 · 被引用 5 次
- Feature Importance Disparities for Data Bias InvestigationsPeter W. Chang, Leor Fishman, Seth NeelICML 2024 · 被引用 3 次
- Evaluating Model Bias Requires Characterizing its MistakesIsabela Albuquerque, Jessica Schrouff, David Warde-Farley, Ali Taylan Cemgil 等ICML 2024 · 被引用 3 次
- Subgroups Matter for Robust Bias MitigationAnissa Alloula, Charles Jones, Ben Glocker, Bartlomiej W. PapiezICML 2025
