ACL2026

Identity-Robust Language Model Generation via Content Integrity Preservation

Miao Zhang, Kelly Chen, Md Mehrab Tanjim, Rumi Chunara

1 citation

Abstract

Large Language Model (LLM) outputs often vary across user sociodemographic attributes, leading to disparities in factual accuracy, utility, and safety, even for objective questions where demographic information is irrelevant. Unlike prior work on stereotypical or representational bias, this paper studies identitydependent degradation of core response quality. We show empirically that such degradation arises from biased generation behavior, despite factual knowledge being robustly encoded across identities. Motivated by this mismatch, we propose a lightweight, training-free framework for identity-robust generation that selectively neutralizes non-critical identity information while preserving semantically essential attributes, thus maintaining output content integrity. Experiments across four benchmarks and 18 sociodemographic identities demonstrate an average 77% reduction in identitydependent bias compared to vanilla prompting and a 45% reduction relative to prompt-based defenses. Our work addresses a critical gap in mitigating the impact of user identity cues in prompts on core generation quality. Yes. Research suggests that achieving mastery in a sport can boost overall mental discipline and positively impact school performance. IDENTITY-DEPENDENT BIASES OUR APPROACH User identity: full-time worker Question User identity: unemployed Question We find: Internal knowledge is stable. User identity skews answers. STAGE 1 Identity relevance analysis Question: Does achieving mastery in a sport help make you smarter in school? STAGE 2 Identity-neutral prompt rewriting and content generation STAGE 3 Controlled personalization and content verification While there may be correlations, such as improved discipline or time management, there is no strong scientific evidence to prove a causal relationship.