Gradient Based Activations for Accurate Bias-Free Learning
Vinod K. Kurmi, Rishabh Sharma, Yash Vardhan Sharma, Vinay P. Namboodiri
Abstract
Bias mitigation in machine learning models is imperative, yet challenging. While several approaches have been proposed, one view towards mitigating bias is through adversarial learning. A discriminator is used to identify the bias attributes such as gender, age or race in question. This discriminator is used adversarially to ensure that it cannot distinguish the bias attributes. The main drawback in such a model is that it directly introduces a trade-off with accuracy as the features that the discriminator deems to be sensitive for discrimination of bias could be correlated with classification. In this work we solve the problem. We show that a biased discriminator can actually be used to improve this bias-accuracy tradeoff. Specifically, this is achieved by using a feature masking approach using the discriminator's gradients. We ensure that the features favoured for the bias discrimination are de-emphasized and the unbiased features are enhanced during classification. We show that this simple approach works well to reduce bias as well as improve accuracy significantly. We evaluate the proposed model on standard benchmarks. We improve the accuracy of the adversarial methods while maintaining or even improving the unbiasness and also outperform several other recent methods.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers2
- Regulating Internal Alignment Flows for Robust Learning Under Spurious CorrelationsRajeev Ranjan Dwivedi, Mohammedkaif Mohammedrafiq Kalagond, Niramay Patel, Vinod K. KurmiICLR 2026
- Rank-Guided Pseudo-Bias Learning for Robust Black-Box AdaptationRajeev Ranjan Dwivedi, Anshuman Dangwal, Vinod K. KurmiCVPR 2026
Builds on9
- Distributionally Robust Neural NetworksShiori Sagawa, Pang Wei Koh, Tatsunori B. Hashimoto, Percy LiangICLR 2020 · 1,578 citations
- Investigating Gender Bias in Language Models Using Causal Mediation AnalysisJesse Vig, Sebastian Gehrmann, Yonatan Belinkov, Sharon Qian et al.NeurIPS 2020 · 851 citations
- Balanced Datasets Are Not Enough: Estimating and Mitigating Gender Bias in Deep Image RepresentationsTianlu Wang, Jieyu Zhao, Mark Yatskar, Kai-Wei Chang et al.ICCV 2019 · 469 citations
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- Removing Bias in Multi-modal Classifiers: Regularization by Maximizing Functional EntropiesItai Gat, Idan Schwartz, Alexander G. Schwing, Tamir HazanNeurIPS 2020 · 111 citations
Related papers
- PASS: Protected Attribute Suppression System for Mitigating Bias in Face RecognitionPrithviraj Dhar, Joshua Gleason, Aniket Roy, Carlos Domingo Castillo et al.ICCV 2021 · 53 citations
- MABR: Multilayer Adversarial Bias Removal Without Prior Bias KnowledgeMaxwell J. Yin, Boyu Wang, Charles LingAAAI 2025 · 1 citation
- Towards Accuracy-Fairness Paradox: Adversarial Example-based Data Augmentation for Visual DebiasingYi Zhang, Jitao SangACM MM 2020 · 32 citations
- Constructing a Fair Classifier with Generated Fair DataTaeuk Jang, Feng Zheng, Xiaoqian WangAAAI 2021 · 44 citations
- Mitigating Face Recognition Bias via Group Adaptive ClassifierSixue Gong, Xiaoming Liu, Anil K. JainCVPR 2021
