Identity-Robust Language Model Generation via Content Integrity Preservation
Miao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi Chunara
Abstract
Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior work on stereotypical or representational bias, this paper studies identitydependent degradation of core response quality. We show empirically that such degradation arises from biased generation behavior, despite factual knowledge being robustly encoded across identities. Motivated by this mismatch, we propose a lightweight, training-free framework for identity-robust generation that selectively neutralizes non-critical identity information while preserving semantically essential attributes, thus maintaining output content integrity. Experiments across four benchmarks and 18 sociodemographic identities demonstrate an average 77% reduction in identitydependent bias compared to vanilla prompting and a 45% reduction relative to prompt-based defenses. Our work addresses a critical gap in mitigating the impact of user identity cues in prompts on core generation quality. Yes. Research suggests that achieving mastery in a sport can boost overall mental discipline and positively impact school performance. IDENTITY-DEPENDENT BIASES OUR APPROACH User identity: full-time worker Question User identity: unemployed Question We find: Internal knowledge is stable. User identity skews answers. STAGE 1 Identity relevance analysis Question: Does achieving mastery in a sport help make you smarter in school? STAGE 2 Identity-neutral prompt rewriting and content generation STAGE 3 Controlled personalization and content verification While there may be correlations, such as improved discipline or time management, there is no strong scientific evidence to prove a causal relationship.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 37765326-a48e-430f-aa78-563240557f26Builds on19
- TruthfulQA: Measuring How Models Mimic Human FalsehoodsStephanie Lin, Jacob Hilton, Owain EvansACL 2022 · 3,228 citations
- AmbigQA: Answering Ambiguous Open-domain QuestionsSewon Min, Julian Michael, Hannaneh Hajishirzi, Luke ZettlemoyerEMNLP 2020 · 162 citations
- Decision-Making Behavior Evaluation Framework for LLMs under Uncertain ContextJingru Jia, Zehua Yuan, Junhao Pan, Paul McNamara et al.NeurIPS 2024 · 71 citations
- Who's asking? User personas and the mechanics of latent misalignmentAsma Ghandeharioun, Ann Yuan, Marius Guerard, Emily Reif et al.NeurIPS 2024 · 44 citations
- UniBias: Unveiling and Mitigating LLM Bias through Internal Attention and FFN ManipulationHanzhang Zhou, Zijian Feng, Zixiao Zhu, Junlang Qian et al.NeurIPS 2024 · 43 citations
Related papers
- Knowing Bias, Doing Better: Mitigating Social Bias in LLMs via Know-Bias Neuron EnhancementJinhao Pan, Chahat Raj, Anjishnu Mukherjee, Sina Mansouri et al.ICML 2026
- LIDAO: Towards Limited Interventions for Debiasing (Large) Language ModelsTianci Liu, Haoyu Wang, Shiyang Wang, Yu Cheng et al.ICML 2024 · 3 citations
- One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM PersonalizationFranziska Weeber, Vera Neplenbroek, Jan Batzner, Sebastian PadóACL 2026 · 4 citations
- Reading Between the Prompts: How Stereotypes Shape LLM's Implicit PersonalizationVera Neplenbroek, Arianna Bisazza, Raquel FernándezEMNLP 2025
- Principled Personas: Defining and Measuring the Intended Effects of Persona Prompting on Task PerformancePedro Henrique Luz de Araujo, Paul Röttger, Dirk Hovy, Benjamin RothEMNLP 2025 · 1 citation
