Adversarial Scrubbing of Demographic Information for Text Classification
Somnath Basu Roy Chowdhury, Sayan Ghosh, Yiyuan Li, Junier Oliva, Shashank Srivastava, Snigdha Chaturvedi
Abstract
Contextual representations learned by language models can often encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated target task. We aim to scrub such undesirable attributes and learn fair representations while maintaining performance on the target task. In this paper, we present an adversarial learning framework "Adversarial Scrubber" (ADS), to debias contextual representations. We perform theoretical analysis to show that our framework converges without leaking demographic information under certain conditions. We extend previous evaluation techniques by evaluating debiasing performance using Minimum Description Length (MDL) probing. Experimental evaluations on 8 datasets show that ADS generates representations with minimal information about demographic attributes while being maximally informative about the target task.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 470cb6ce-a143-4a88-802a-0333331e6d20Cited by top-tier papers7
- Robust Concept Erasure via Kernelized Rate-Distortion MaximizationSomnath Basu Roy Chowdhury, Nicholas Monath, Kumar Avinava Dubey, Amr Ahmed et al.NeurIPS 2023 · 12 citations
- Dual-Teacher De-Biasing Distillation Framework for Multi-Domain Fake News DetectionJiayang Li, Xuan Feng, Tianlong Gu, Liang ChangICDE 2024 · 10 citations
- Sustaining Fairness via Incremental LearningSomnath Basu Roy Chowdhury, Snigdha ChaturvediAAAI 2023 · 6 citations
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett et al.EMNLP 2024 · 4 citations
- Enhancing Group Fairness in Online Settings Using Oblique Decision ForestsSomnath Basu Roy Chowdhury, Nicholas Monath, Ahmad Beirami, Rahul Kidambi et al.ICLR 2024 · 3 citations
Builds on6
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 93 citations
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 68 citations
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial LearningHaochen Liu, Wentao Wang, Yiqi Wang, Hui Liu et al.EMNLP 2020 · 55 citations
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 34 citations
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston et al.EMNLP 2020 · 7 citations
Related papers
- Everybody Needs Good Neighbours: An Unsupervised Locality-based Method for Bias MitigationXudong Han, Timothy Baldwin, Trevor CohnICLR 2023
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 4 citations
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial InferenceYuhong Luo, Austin Hoag, Xintong Wang, Philip S. Thomas et al.NeurIPS 2025
- Prompt Tuning Pushes Farther, Contrastive Learning Pulls Closer: A Two-Stage Approach to Mitigate Social BiasesYingji Li, Mengnan Du, Xin Wang, Ying WangACL 2023 · 12 citations
- LIDAO: Towards Limited Interventions for Debiasing (Large) Language ModelsTianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng et al.ICML 2024 · 3 citations
