Mind the Graph When Balancing Data for Fairness or Robustness
Jessica Schrouff, Alexis Bellot, Amal Rannen-Triki, Alan Malek, Isabela Albuquerque, Arthur Gretton, Alexander D'Amour, Silvia Chiappa
Abstract
Failures of fairness or robustness in machine learning predictive settings can be due to undesired dependencies between covariates, outcomes and auxiliary factors of variation. A common strategy to mitigate these failures is data balancing, which attempts to remove those undesired dependencies. In this work, we define conditions on the training distribution for data balancing to lead to fair or robust models. Our results display that, in many cases, the balanced distribution does not correspond to selectively removing the undesired dependencies in a causal graph of the task, leading to multiple failure modes and even interference with other mitigation techniques such as regularization. Overall, our results highlight the importance of taking the causal graph into account before performing data balancing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext bf6bf362-17c4-416d-adb6-ba164f01369fCited by top-tier papers5
- Person-Centric Annotations of LAION-400M: Auditing Bias and Its Transfer to ModelsLeander Girrbach, Stephan Alaniz, Genevieve Smith, Trevor Darrell et al.ICLR 2026 · 5 citations
- Does Weak-to-strong Generalization Happen under Spurious Correlations?Chenruo Liu, Yijun Dong, Qi LeiICLR 2026 · 1 citation
- Subgroups Matter for Robust Bias MitigationAnissa Alloula, Charles Jones, Ben Glocker, Bartlomiej W. PapiezICML 2025
- Beyond Binary Erasure: Soft-Weighted Unlearning for Fairness and RobustnessXinbao Qiao, Ningning Ding, Yushi Cheng, Meng ZhangAAAI 2026
- Causal Fine-Tuning under Latent Confounded ShiftJialin Yu, Yuxiang Zhou, Haoxuan Li, Junchi Yu et al.ICML 2026
Builds on25
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
Related papers
- The Importance of Modeling Data Missingness in Algorithmic Fairness: A Causal PerspectiveNaman Goel, Alfonso Amayuelas, Amit Deshpande, Amit SharmaAAAI 2021 · 36 citations
- On Disentangled Representations Learned from Correlated DataFrederik Träuble, Elliot Creager, Niki Kilbertus, Francesco Locatello et al.ICML 2021 · 38 citations
- Correcting Overparameterization Effects in Fair Empirical Risk MinimizationXiaoyi MAI, Jean-Michel LoubesICML 2026
- The Fairness Hierarchy: A viewpoint from causal inferenceChengbo Zhang, Zhen Yao, Hao Pang, Changcheng LiICML 2026
- Towards Robust Classification Model by Counterfactual and Invariant Data GenerationChun-Hao Chang, George-Alexandru Adam, Anna GoldenbergCVPR 2021
