BLIND: Bias Removal With No Demographics
Hadas Orgad, Yonatan Belinkov
Abstract
Models trained on real-world data tend to imitate and amplify social biases. Common methods to mitigate biases require prior information on the types of biases that should be mitigated (e.g., gender or racial bias) and the social groups associated with each data sample. In this work, we introduce BLIND, a method for bias removal with no prior knowledge of the demographics in the dataset. While training a model on a downstream task, BLIND detects biased samples using an auxiliary model that predicts the main model's success, and downweights those samples during the training process. Experiments with racial and gender biases in sentiment classification and occupation classification tasks demonstrate that BLIND mitigates social biases without relying on a costly demographic annotation process. Our method is competitive with other methods that require demographic information and sometimes even surpasses them. 1 * Supported by the Viterbi Fellowship in the Center for Computer Engineering at the Technion. 1 Our code is available at https://github.com/ technion-cs-nlp/BLIND .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext c238562f-a3e8-4954-9988-5784a032ddf3Cited by top-tier papers5
- LLM Whisperer: An Inconspicuous Attack to Bias LLM ResponsesWeiran Lin, Anna Gerchanovsky, Omer Akgul, Lujo Bauer et al.CHI 2025 · 23 citations
- MABR: Multilayer Adversarial Bias Removal Without Prior Bias KnowledgeMaxwell J. Yin, Boyu Wang, Charles LingAAAI 2025 · 1 citation
- Attention Pruning: Automated Fairness Repair of Language Models via Surrogate Simulated AnnealingVishnu Asutosh Dasu, Md Rafi Ur Rashid, Vipul Gupta, Saeid Tizpaz-Niari et al.ICSE 2026
- Sparse Feature Circuits: Discovering and Editing Interpretable Causal Graphs in Language ModelsSamuel Marks, Can Rager, Eric J. Michaud, Yonatan Belinkov et al.ICLR 2025
- Why Don't Prompt-Based Fairness Metrics Correlate?Abdelrahman Zayed, Gonçalo Mordido, Ioana Baldini, Sarath ChandarACL 2024
Builds on15
- Deberta: decoding-Enhanced Bert with Disentangled AttentionPengcheng He, Xiaodong Liu, Jianfeng Gao, Weizhu ChenICLR 2021 · 3,729 citations
- Just Train Twice: Improving Group Robustness without Training Group InformationEvan Zheran Liu, Behzad Haghgoo, Annie S. Chen, Aditi Raghunathan et al.ICML 2021 · 683 citations
- Fairness without Demographics through Adversarially Reweighted LearningPreethi Lahoti, Alex Beutel, Jilin Chen, Kang Lee et al.NeurIPS 2020 · 406 citations
- FairBatch: Batch Selection for Model FairnessYuji Roh, Kangwook Lee, Steven Euijong Whang, Changho SuhICLR 2021 · 156 citations
- End-to-End Bias Mitigation by Modelling Biases in CorporaRabeeh Karimi Mahabadi, Yonatan Belinkov, James HendersonACL 2020 · 136 citations
Related papers
- Interpretable Debiasing of Vision-Language Models for Social FairnessNa Min An, Yoonna Jang, Yusuke Hirota, Ryo Hachiuma et al.CVPR 2026 · 7 citations
- Auto-Debias: Debiasing Masked Language Models with Automated Biased PromptsYue Guo, Yi Yang, Ahmed AbbasiACL 2022
- Towards Debiasing NLU Models from Unknown BiasesPrasetya Ajie Utama, Nafise Sadat Moosavi, Iryna GurevychEMNLP 2020 · 3 citations
- Bias Mimicking: A Simple Sampling Approach for Bias MitigationMaan Qraitem, Kate Saenko, Bryan A. PlummerCVPR 2023
- Causal-Debias: Unifying Debiasing in Pretrained Language Models and Fine-tuning via Causal Invariant LearningFan Zhou, Yuzhou Mao, Liu Yu, Yi Yang et al.ACL 2023 · 21 citations
