Mind the Graph When Balancing Data for Fairness or Robustness
Jessica Schrouff, Alexis Bellot, Amal Rannen-Triki, Alan Malek, Isabela Albuquerque, Arthur Gretton, Alexander D'Amour, Silvia Chiappa
摘要
Failures of fairness or robustness in machine learning predictive settings can be due to undesired dependencies between covariates, outcomes and auxiliary factors of variation. A common strategy to mitigate these failures is data balancing, which attempts to remove those undesired dependencies. In this work, we define conditions on the training distribution for data balancing to lead to fair or robust models. Our results display that, in many cases, the balanced distribution does not correspond to selectively removing the undesired dependencies in a causal graph of the task, leading to multiple failure modes and even interference with other mitigation techniques such as regularization. Overall, our results highlight the importance of taking the causal graph into account before performing data balancing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to ModelsLeander Girrbach, Stephan Alaniz, Genevieve Smith, Trevor Darrell 等ICLR 2026 · 被引用 5 次
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 被引用 1 次
- Subgroups Matter for Robust Bias MitigationAnissa Alloula, Charles Jones, Ben Glocker, Bartlomiej W. PapiezICML 2025
- Beyond Binary Erasure: Soft-Weighted Unlearning for Fairness and RobustnessXinbao Qiao, Ningning Ding, Yushi Cheng, Meng ZhangAAAI 2026
- Causal Fine-Tuning under Latent Confounded ShiftJialin Yu, Yuxiang Zhou, Haoxuan Li, Junchi Yu 等ICML 2026
它引用的顶会 Paper25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa 等ICML 2021 · 被引用 8,974 次
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 被引用 1,578 次
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang 等ICCV 2019 · 被引用 469 次
相关 Paper
- The Importance of Modeling Data Missingness in Algorithmic Fairness: A Causal PerspectiveNaman Goel, Alfonso Amayuelas, Amit Deshpande, Amit SharmaAAAI 2021 · 被引用 36 次
- On Disentangled Representations Learned from Correlated DataFrederik Träuble, Elliot Creager, Niki Kilbertus, Francesco Locatello 等ICML 2021 · 被引用 38 次
- Correcting Overparameterization Effects in Fair Empirical Risk MinimizationXiaoyi MAI, Jean-Michel LoubesICML 2026
- The Fairness Hierarchy: A viewpoint from causal inferenceChengbo Zhang, Zhen Yao, Hao Pang, Changcheng LiICML 2026
- Towards Robust Classification Model by Counterfactual and Invariant Data GenerationChun-Hao Chang, George-Alexandru Adam, Anna GoldenbergCVPR 2021
