Adversarial Scrubbing of Demographic Information for Text Classification
Somnath Basu Roy Chowdhury, Sayan Ghosh, Yiyuan Li, Junier Oliva, Shashank Srivastava, Snigdha Chaturvedi
摘要
Contextual representations learned by language models can often encode undesirable attributes, like demographic associations of the users, while being trained for an unrelated target task. We aim to scrub such undesirable attributes and learn fair representations while maintaining performance on the target task. In this paper, we present an adversarial learning framework "Adversarial Scrubber" (ADS), to debias contextual representations. We perform theoretical analysis to show that our framework converges without leaking demographic information under certain conditions. We extend previous evaluation techniques by evaluating debiasing performance using Minimum Description Length (MDL) probing. Experimental evaluations on 8 datasets show that ADS generates representations with minimal information about demographic attributes while being maximally informative about the target task.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper7
- Robust Concept Erasure via Kernelized Rate-Distortion MaximizationSomnath Basu Roy Chowdhury, Nicholas Monath, Kumar Avinava Dubey, Amr Ahmed 等NeurIPS 2023 · 被引用 12 次
- Dual-Teacher De-Biasing Distillation Framework for Multi-Domain Fake News DetectionJiayang Li, Xuan Feng, Tianlong Gu, Liang ChangICDE 2024 · 被引用 10 次
- Sustaining Fairness via Incremental LearningSomnath Basu Roy Chowdhury, Snigdha ChaturvediAAAI 2023 · 被引用 6 次
- Metrics for What, Metrics for Whom: Assessing Actionability of Bias Evaluation Metrics in NLPPieter Delobelle, Giuseppe Attanasio, Debora Nozza, Su Lin Blodgett 等EMNLP 2024 · 被引用 4 次
- Enhancing Group Fairness in Online Settings Using Oblique Decision ForestsSomnath Basu Roy Chowdhury, Nicholas Monath, Ahmad Beirami, Rahul Kidambi 等ICLR 2024 · 被引用 3 次
它引用的顶会 Paper6
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 被引用 93 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial LearningHaochen Liu, Wentao Wang, Yiqi Wang, Hui Liu 等EMNLP 2020 · 被引用 55 次
- Information-Theoretic Probing with Minimum Description LengthElena Voita, Ivan TitovEMNLP 2020 · 被引用 34 次
- Multi-Dimensional Gender Bias ClassificationEmily Dinan, Angela Fan, Ledell Wu, Jason Weston 等EMNLP 2020 · 被引用 7 次
相关 Paper
- Everybody Needs Good Neighbours: An Unsupervised Locality-based Method for Bias MitigationXudong Han, Timothy Baldwin, Trevor CohnICLR 2023
- Mitigating Biases in Language Models via Bias UnlearningDianqing Liu, Yi Liu, Guoqing Jin, Zhendong MaoEMNLP 2025 · 被引用 4 次
- Fair Representation Learning with Controllable High Confidence Guarantees via Adversarial InferenceYuhong Luo, Austin Hoag, Xintong Wang, Philip S. Thomas 等NeurIPS 2025
- Prompt Tuning Pushes Farther, Contrastive Learning Pulls Closer: A Two-Stage Approach to Mitigate Social BiasesYingji Li, Mengnan Du, Xin Wang, Ying WangACL 2023 · 被引用 12 次
- LIDAO: Towards Limited Interventions for Debiasing (Large) Language ModelsTianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng 等ICML 2024 · 被引用 3 次
