ChatGPT Doesn't Trust Chargers Fans: Guardrail Sensitivity in Context
Victoria R. Li, Yida Chen, Naomi Saphra
Abstract
While the biases of language models in production are extensively documented, the biases of their guardrails have been neglected. This paper studies how contextual information about the user influences the likelihood of an LLM to refuse to execute a request. By generating user biographies that offer ideological and demographic information, we find a number of biases in guardrail sensitivity on GPT-3.5. Younger, female, and Asian-American personas are more likely to trigger a refusal guardrail when requesting censored or illegal information. Guardrails are also sycophantic, refusing to comply with requests for a political position the user is likely to disagree with. We find that certain identity groups and seemingly innocuous information, e.g., sports fandom, can elicit changes in guardrail sensitivity similar to direct statements of political ideology. For each demographic category and even for American football team fandom, we find that ChatGPT appears to infer a likely political ideology and modify guardrail behavior accordingly.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers4
- The Alignment Waltz: Jointly Training Agents to Collaborate for SafetyJingyu Zhang, Haozhu Wang, Eric Michael Smith, Sid Wang et al.ICLR 2026 · 11 citations
- Does Refusal Training in LLMs Generalize to the Past Tense?Maksym Andriushchenko, Nicolas FlammarionICLR 2025 · 6 citations
- Creating and Evaluating Personas Using Generative AI: A Scoping Review of 81 ArticlesDanial Amin, Joni Salminen, Farhan Ahmed, Sonja M. H. Tervola et al.CHI 2026 · 6 citations
- Investigating Counterfactual Unfairness in LLMs towards Identities through HumorShubin Kim, Yejin Son, Junyeong Park, Keummin Ka et al.ACL 2026
Builds on3
- Fine-tuning Aligned Language Models Compromises Safety, Even When Users Do Not Intend To!Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen et al.ICLR 2024 · 1,104 citations
- Whose Opinions Do Language Models Reflect?Shibani Santurkar, Esin Durmus, Faisal Ladhak, Cinoo Lee et al.ICML 2023 · 764 citations
- Evaluation of African American Language Bias in Natural Language GenerationNicholas Deas, Jessica Grieser, Shana Kleiner, Desmond Patton et al.EMNLP 2023 · 15 citations
Related papers
- Reading Between the Prompts: How Stereotypes Shape LLM's Implicit PersonalizationVera Neplenbroek, Arianna Bisazza, Raquel FernándezEMNLP 2025
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsShashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan et al.ICLR 2024 · 212 citations
- One Persona, Many Cues, Different Results: How Sociodemographic Cues Impact LLM PersonalizationFranziska Weeber, Vera Neplenbroek, Jan Batzner, Sebastian PadóACL 2026 · 4 citations
- There is No War in Ba Sing Se: A Global Analysis of Content Moderation in Large Language ModelsFriedemann Lipphardt, Moonis Ali, Martin Banzer, Anja Feldmann et al.NDSS 2026 · 1 citation
- RST-Guarder: Enhancing Long-Context Robustness for Safeguards via RST Parsing and Probabilistic InferenceXu Zhang, Xiaojun WanACL 2026
