Identity-Robust Language Model Generation via Content Integrity Preservation
Miao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi Chunara
摘要
Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior work on stereotypical or representational bias, this paper studies identitydependent degradation of core response quality. We show empirically that such degradation arises from biased generation behavior, despite factual knowledge being robustly encoded across identities. Motivated by this mismatch, we propose a lightweight, training-free framework for identity-robust generation that selectively neutralizes non-critical identity information while preserving semantically essential attributes, thus maintaining output content integrity. Experiments across four benchmarks and 18 sociodemographic identities demonstrate an average 77% reduction in identitydependent bias compared to vanilla prompting and a 45% reduction relative to prompt-based defenses. Our work addresses a critical gap in mitigating the impact of user identity cues in prompts on core generation quality. Yes. Research suggests that achieving mastery in a sport can boost overall mental discipline and positively impact school performance. IDENTITY-DEPENDENT BIASES OUR APPROACH User identity: full-time worker Question User identity: unemployed Question We find: Internal knowledge is stable. User identity skews answers. STAGE 1 Identity relevance analysis Question: Does achieving mastery in a sport help make you smarter in school? STAGE 2 Identity-neutral prompt rewriting and content generation STAGE 3 Controlled personalization and content verification While there may be correlations, such as improved discipline or time management, there is no strong scientific evidence to prove a causal relationship.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper19
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 被引用 3,228 次
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 被引用 162 次
- Decision-Making Behavior Evaluation Framework for LLMs under Uncertain ContextJingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara 等NeurIPS 2024 · 被引用 71 次
- Who's asking? User personas and the mechanics of latent misalignmentAsma Ghandeharioun, Ann Yuan, Marius Guerard, Emily Reif 等NeurIPS 2024 · 被引用 44 次
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN ManipulationHanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian 等NeurIPS 2024 · 被引用 43 次
相关 Paper
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri 等ICML 2026
- LIDAO: Towards Limited Interventions for Debiasing (Large) Language ModelsTianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng 等ICML 2024 · 被引用 3 次
- One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM PersonalizationFranziska Weeber, Vera Neplenbroek, Jan Batzner, Sebastian PadóACL 2026 · 被引用 4 次
- Reading Between the Prompts: How Stereotypes Shape LLM's Implicit PersonalizationVera Neplenbroek, Arianna Bisazza, Raquel FernándezEMNLP 2025
- Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task PerformancePedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin RothEMNLP 2025 · 被引用 1 次
