OSCaR: Orthogonal Subspace Correction and Rectification of Biases in Word Embeddings
Sunipa Dev, Tao Li, Jeff M. Phillips, Vivek Srikumar
摘要
Language representations are known to carry certain associations (e.g., gendered connotations) which may lead to invalid and harmful predictions in downstream tasks. While existing methods are effective at mitigating such unwanted associations by linear projection, we argue that they are too aggressive: not only do they remove such associations, they also erase information that should be retained. To address this issue, we propose OS-CAR (Orthogonal Subspace Correction and Rectification), a balanced approach of mitigation that focuses on disentangling associations between concepts that are deemed problematic, instead of removing concepts wholesale. We develop new measurements for evaluating information retention relevant to the debiasing goal. Our experiments on genderoccupation associations show that OSCAR is a well-balanced approach that ensures that semantic information is retained in the embeddings and unwanted associations are also effectively mitigated. Quantifying Bias and Information Retained Keeping in mind that removing bias and retaining information have to be done in synergy, we present how to obtain aggregated measurements for these two components. We will first describe the de-
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper16
- LEACE: Perfect linear concept erasure in closed formNora Belrose, David Schneider-Joseph, Shauli Ravfogel, Ryan Cotterell 等NeurIPS 2023 · 被引用 305 次
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian 等EMNLP 2021 · 被引用 113 次
- Linear Adversarial Concept ErasureShauli Ravfogel, Michael Twiton, Yoav Goldberg, Ryan CotterellICML 2022 · 被引用 89 次
- MABEL: Attenuating Gender Bias using Textual Entailment DataJacqueline He, Mengzhou Xia, Christiane Fellbaum, Danqi ChenEMNLP 2022 · 被引用 15 次
- SeeGULL: A Stereotype Benchmark with Broad Geo-Cultural Coverage Leveraging Generative ModelsAkshita Jha, Aida Mostafazadeh Davani, Chandan K. Reddy, Shachi Dave 等ACL 2023 · 被引用 13 次
它引用的顶会 Paper7
- On Measuring and Mitigating Biased Inferences of Word EmbeddingsSunipa Dev, Tao Li, Jeff M. Phillips, Vivek SrikumarAAAI 2020 · 被引用 195 次
- Harms of Gender Exclusivity and Challenges in Non-Binary Representation in Language TechnologiesSunipa Dev, Masoud Monajatipoor, Anaelia Ovalle, Arjun Subramonian 等EMNLP 2021 · 被引用 113 次
- Language (Technology) is Power: A Critical Survey of "Bias" in NLPSu Lin Blodgett, Solon Barocas, Hal Daumé III, Hanna M. WallachACL 2020 · 被引用 68 次
- A General Framework for Implicit and Explicit Debiasing of Distributional Word Vector SpacesAnne Lauscher, Goran Glavas, Simone Paolo Ponzetto, Ivan VulicAAAI 2020 · 被引用 68 次
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
相关 Paper
- Towards Debiasing Sentence RepresentationsPaul Pu Liang, Irene Mengze Li, Emily Zheng, Yao Chong Lim 等ACL 2020 · 被引用 149 次
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 被引用 4 次
- Bias Is a Subspace, Not a Coordinate: A Geometric Rethinking of Post‑hoc Debiasing in Vision-Language ModelsDachuan Zhao, Weiyue Li, Zhenda Shen, Yushu Qiu 等CVPR 2026 · 被引用 5 次
- Conceptor-Aided Debiasing of Large Language ModelsYifei Li, Lyle H. Ungar, João SedocEMNLP 2023 · 被引用 5 次
- Exploring the Linear Subspace Hypothesis in Gender Bias MitigationFrancisco Vargas, Ryan CotterellEMNLP 2020 · 被引用 2 次
