Towards Debiasing DNN Models from Spurious Feature Influence
Mengnan Du, Ruixiang Tang, Weijie Fu, Xia Hu
摘要
Recent studies indicate that deep neural networks (DNNs) are prone to show discrimination towards certain demographic groups. We observe that algorithmic discrimination can be explained by the high reliance of the models on fairness sensitive features. Motivated by this observation, we propose to achieve fairness by suppressing the DNN models from capturing the spurious correlation between those fairness sensitive features with the underlying task. Specifically, we firstly train a bias-only teacher model which is explicitly encouraged to maximally employ fairness sensitive features for prediction. The teacher model then counter-teaches a debiased student model so that the interpretation of the student model is orthogonal to the interpretation of the teacher model. The key idea is that since the teacher model relies explicitly on fairness sensitive features for prediction, the orthogonal interpretation loss enforces the student network to reduce its reliance on sensitive features and instead capture more task relevant features for prediction. Experimental analysis indicates that our framework substantially reduces the model's attention on fairness sensitive features. Experimental results on four datasets further validate that our framework has consistently improved the fairness with respect to three group fairness metrics, with a comparable or even better accuracy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper1
相关 Paper
- Information-Theoretic Bias Reduction via Causal View of Spurious CorrelationSeonguk Seo, Joon-Young Lee, Bohyung HanAAAI 2022 · 被引用 29 次
- Controllable Feature Whitening for Hyperparameter-Free Bias MitigationYooshin Cho, Hanbyel Cho, Janghyeon Lee, Hyeong Gwon Hong 等ICCV 2025 · 被引用 2 次
- Let Samples Speak: Mitigating Spurious Correlation by Exploiting the Clusterness of SamplesWeiwei Li, Junzhuo Liu, Yuanyuan Ren, Yuchen Zheng 等CVPR 2025
- Fairness via Representation NeutralizationMengnan Du, Subhabrata Mukherjee, Guanchu Wang, Ruixiang Tang 等NeurIPS 2021 · 被引用 91 次
- Benign Shortcut for Debiasing: Fair Visual Recognition via Intervention with Shortcut FeaturesYi Zhang, Jitao Sang, Junyang Wang, Dongmei Jiang 等ACM MM 2023 · 被引用 9 次
