Counterfactually Fair Representation
Zhiqun Zuo, Mahdi Khalili, Xueru Zhang
Abstract
The use of machine learning models in high-stake applications (e.g., healthcare, lending, college admission) has raised growing concerns due to potential biases against protected social groups. Various fairness notions and methods have been proposed to mitigate such biases. In this work, we focus on Counterfactual Fairness (CF), a fairness notion that is dependent on an underlying causal graph and first proposed by Kusner et al. ; it requires that the outcome an individual perceives is the same in the real world as it would be in a"counterfactual"world, in which the individual belongs to another social group. Learning fair models satisfying CF can be challenging. It was shown in that a sufficient condition for satisfying CF is to not use features that are descendants of sensitive attributes in the causal graph. This implies a simple method that learns CF models only using non-descendants of sensitive attributes while eliminating all descendants. Although several subsequent works proposed methods that use all features for training CF models, there is no theoretical guarantee that they can satisfy CF. In contrast, this work proposes a new algorithm that trains models using all the available features. We theoretically and empirically show that models trained with this method can satisfy CF.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers5
- Counterfactual Fairness by Combining Factual and Counterfactual PredictionsZeyu Zhou, Tianci Liu, Ruqi Bai, Jing Gao et al.NeurIPS 2024 · 11 citations
- Learning Counterfactual Outcomes Under Rank PreservationPeng Wu, Haoxuan Li, Chunyuan Zheng, Yan Zeng et al.NeurIPS 2025 · 7 citations
- Benchmarking Bias Mitigation Toward Fairness Without Harm from Vision to LVLMsXuwei Tan, Ziyu Hu, Xueru ZhangICLR 2026 · 4 citations
- CARL: Preserving Causal Structure in Representation LearningYulong Li, Xiwei Liu, Feilong Tang, Zhixiang Lu et al.ICLR 2026
- Individual Fairness In Strategic ClassificationZhiqun Zuo, Mohammad Mahdi KhaliliNeurIPS 2025
Builds on9
- Towards Personalized Fairness based on Causal NotionYunqi Li, Hanxiong Chen, Shuyuan Xu, Yingqiang Ge et al.SIGIR 2021 · 139 citations
- Post-processing for Individual FairnessFelix Petersen, Debarghya Mukherjee, Yuekai Sun, Mikhail YurochkinNeurIPS 2021 · 115 citations
- Counterfactual Fairness with Disentangled Causal Effect Variational AutoencoderHyemi Kim, Seungjae Shin, JoonHo Jang, Kyungwoo Song et al.AAAI 2021 · 72 citations
- Online Certification of Preference-Based Fairness for Personalized Recommender SystemsVirginie Do, Sam Corbett-Davies, Jamal Atif, Nicolas UsunierAAAI 2022 · 47 citations
- Fairness Interventions as (Dis)Incentives for Strategic ManipulationXueru Zhang, Mohammad Mahdi Khalili, Kun Jin, Parinaz Naghizadeh et al.ICML 2022 · 27 citations
Related papers
- Learning for Counterfactual Fairness from Observational DataJing Ma, Ruocheng Guo, Aidong Zhang, Jundong LiKDD 2023 · 9 citations
- Counterfactual Fairness with Partially Known Causal GraphAoqi Zuo, Susan Wei, Tongliang Liu, Bo Han et al.NeurIPS 2022 · 32 citations
- The Fairness Hierarchy: A viewpoint from causal inferenceChengbo Zhang, Zhen Yao, Hao Pang, Changcheng LiICML 2026
- Counterfactual Fairness with Imperfect Causal GraphsCong Su, Qiaoyu Tan, Carlotta Domeniconi, Lizhen Cui et al.AAAI 2026
- Counterfactual Fairness Through Transforming Data Orthogonal to BiasShuyi Chen, Shixiang ZhuKDD 2025
