Unsupervised Concept Vector Extraction for Bias Control in LLMs
Hannah Cyberey, Yangfeng Ji, David Evans
摘要
Large language models (LLMs) are known to perpetuate stereotypes and exhibit biases. Various strategies have been proposed to mitigate these biases, but most work studies biases as a black-box problem without considering how concepts are represented within the model. We adapt techniques from representation engineering to study how the concept of"gender"is represented within LLMs. We introduce a new method that extracts concept representations via probability weighting without labeled data and efficiently selects a steering vector for measuring and manipulating the model's representation. We develop a projection-based method that enables precise steering of model predictions and demonstrate its effectiveness in mitigating gender bias in LLMs and show that it also generalizes to racial bias. Our code is available at: https://github.com/hannahxchen/gender-bias-steering
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper9
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Refusal in Language Models Is Mediated by a Single DirectionAndy Arditi, Oscar Obeso, Aaquib Syed, Daniel Paleka 等NeurIPS 2024 · 被引用 1,166 次
- Language Models Represent Space and TimeWes Gurnee, Max TegmarkICLR 2024 · 被引用 303 次
- Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language ModelsMyra Cheng, Esin Durmus, Dan JurafskyACL 2023 · 被引用 89 次
- Detecting Gender Stereotypes: Lexicon vs. Supervised Learning MethodsJenna Cryan, Shiliang Tang, Xinyi Zhang, Miriam J. Metzger 等CHI 2020 · 被引用 43 次
相关 Paper
- Bi-directional Bias Attribution: Debiasing Large Language Models without Modifying PromptsYujie Lin, Kunquan Li, Yixuan Liao, Xiaoxin Chen 等ICLR 2026 · 被引用 6 次
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 被引用 495 次
- Debiasing Algorithm through Model AdaptationTomasz Limisiewicz, David Marecek, Tomás MusilICLR 2024 · 被引用 24 次
- Fairness Mediator: Neutralize Stereotype Associations to Mitigate Bias in Large Language ModelsYisong Xiao, Aishan Liu, Siyuan Liang, Xianglong Liu 等ISSTA 2025 · 被引用 2 次
- On Optimal Steering to Achieve Exact FairnessMohit Sharma, Amit Deshpande, Chiranjib Bhattacharyya, Rajiv Ratn ShahNeurIPS 2025 · 被引用 2 次
