Marked Personas: Using Natural Language Prompts to Measure Stereotypes in Language Models
Myra Cheng, Esin Durmus, Dan Jurafsky
Abstract
To recognize and mitigate harms from large language models (LLMs), we need to understand the prevalence and nuances of stereotypes in LLM outputs. Toward this end, we present Marked Personas, a prompt-based method to measure stereotypes in LLMs for intersectional demographic groups without any lexicon or data labeling. Grounded in the sociolinguistic concept of markedness (which characterizes explicitly linguistically marked categories versus unmarked defaults), our proposed method is twofold: 1) prompting an LLM to generate personas, i.e., natural language descriptions, of the target demographic group alongside personas of unmarked, default groups; 2) identifying the words that significantly distinguish personas of the target group from corresponding unmarked ones. We find that the portrayals generated by GPT-3.5 and GPT-4 contain higher rates of racial stereotypes than human-written portrayals using the same prompts. The words distinguishing personas of marked (non-white, non-male) groups reflect patterns of othering and exoticizing these demographics. An intersectional lens further reveals tropes that dominate portrayals of marginalized groups, such as tropicalism and the hypersexualization of minoritized women. These representational harms have concerning implications for downstream applications like story generation.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d024bc05-ae0c-4984-9acd-b598850e720bCited by top-tier papers51
- SOTOPIA: Interactive Evaluation for Social Intelligence in Language AgentsXuhui Zhou, Hao Zhu, Leena Mathur, Ruohong Zhang et al.ICLR 2024 · 288 citations
- Bias Runs Deep: Implicit Reasoning Biases in Persona-Assigned LLMsShashank Gupta, Vaishnavi Shrivastava, Ameet Deshpande, Ashwin Kalyan et al.ICLR 2024 · 212 citations
- Apathetic or Empathetic? Evaluating LLMs' Emotional Alignments with HumansJen-tse Huang, Man Ho Lam, Eric John Li, Shujie Ren et al.NeurIPS 2024 · 63 citations
- Rehearsal: Simulating Conflict to Teach Conflict ResolutionOmar Shaikh, Valentino Emil Chai, Michele Gelfand, Diyi Yang et al.CHI 2024 · 62 citations
- Deus Ex Machina and Personas from Large Language Models: Investigating the Composition of AI-Generated Persona DescriptionsJoni Salminen, Chang Liu, Wenjing Pian, Jianxing Chi et al.CHI 2024 · 55 citations
Builds on10
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- CoAuthor: Designing a Human-AI Collaborative Writing Dataset for Exploring Language Model CapabilitiesMina Lee, Percy Liang, Qian YangCHI 2022 · 340 citations
- Bias Out-of-the-Box: An Empirical Analysis of Intersectional Occupational Biases in Popular Generative Language ModelsHannah Rose Kirk, Yennie Jun, Filippo Volpin, Haider Iqbal et al.NeurIPS 2021 · 243 citations
- "I'm sorry to hear that": Finding New Biases in Language Models with a Holistic Descriptor DatasetEric Michael Smith, Melissa Hall, Melanie Kambadur, Eleonora Presani et al.EMNLP 2022 · 56 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
Related papers
- Reading Between the Prompts: How Stereotypes Shape LLM's Implicit PersonalizationVera Neplenbroek, Arianna Bisazza, Raquel FernándezEMNLP 2025
- Who is better at math, Jenny or Jingzhen? Uncovering Stereotypes in Large Language ModelsZara Siddique, Liam D. Turner, Luis Espinosa AnkeEMNLP 2024 · 2 citations
- Towards Understanding and Mitigating Social Biases in Language ModelsPaul Pu Liang, Chiyu Wu, Louis-Philippe Morency, Ruslan SalakhutdinovICML 2021 · 495 citations
- More of the Same: Persistent Representational Harms Under Increased RepresentationJennifer Mickel, Maria De-Arteaga, Liu Leqi, Kevin TianNeurIPS 2025 · 9 citations
- CoMPosT: Characterizing and Evaluating Caricature in LLM SimulationsMyra Cheng, Tiziano Piccardi, Diyi YangEMNLP 2023 · 33 citations
