Interpretable Debiasing of Vectorized Language Representations with Iterative Orthogonalization
Prince Osei Aboagye, Yan Zheng, Jack Shunn, Chin-Chia Michael Yeh, Junpeng Wang, Zhongfang Zhuang, Huiyuan Chen, Liang Wang, Wei Zhang, Jeff M. Phillips
摘要
We propose a new mechanism to augment a word vector embedding representation that offers improved bias removal while retaining the key information—resulting in improved interpretability of the representation. Rather than removing the information associated with a concept that may induce bias, our proposed method identifies two concept subspaces and makes them orthogonal. The resulting representation has these two concepts uncorrelated. Moreover, because they are orthogonal, one can simply apply a rotation on the basis of the representation so that the resulting subspace corresponds with coordinates. This explicit encoding of concepts to coordinates works because they have been made fully orthogonal, which previous approaches do not achieve. Furthermore, we show that this can be extended to multiple subspaces. As a result, one can choose a subset of concepts to be represented transparently and explicitly, while the others are retained in the mixed but extremely expressive format of the representation.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
引用它的顶会 Paper2
- Zero-Shot Robustification of Zero-Shot ModelsDyah Adila, Changho Shin, Linrong Cai, Frederic SalaICLR 2024 · 被引用 31 次
- Model Editing as a Robust and Denoised variant of DPO: A Case Study on ToxicityRheeya Uppaal, Apratim Dey, Yiting He, Yiqiao Zhong 等ICLR 2025
相关 Paper
- WRING Out The Bias: A Rotation-Based Alternative To Projection DebiasingWalter Gerych, Cassandra Parent, Quinn Perian, Rafiya Javed 等ICLR 2026
- Null It Out: Guarding Protected Attributes by Iterative Nullspace ProjectionShauli Ravfogel, Yanai Elazar, Hila Gonen, Michael Twiton 等ACL 2020 · 被引用 25 次
- Preserving Task-Relevant Information Under Linear Concept RemovalFloris Holstege, Shauli Ravfogel, Bram WoutersNeurIPS 2025 · 被引用 4 次
- Removing Spurious Concepts from Neural Network Representations via Joint Subspace EstimationFloris Holstege, Bram Wouters, Noud P. A. van Giersbergen, Cees G. H. DiksICML 2024 · 被引用 3 次
- Decomposing Representation Space into Interpretable Subspaces with Unsupervised LearningXinting Huang, Michael HahnICLR 2026 · 被引用 7 次
