Dark and Bright Side of Participatory Red-Teaming with Targets of Stereotyping for Eliciting Harmful Behaviors from Large Language Models
Sieun Kim, Yeeun Jo, Sungmin Na, Hyunseung Lim, Eunchae Lee, Yu Min Choi, Soohyun Cho, Hwajung Hong
Abstract
Warning: This article contains stereotypical and offensive content.
Red-teaming-where adversarial prompts are crafted to expose harmful behaviors and assess risks-offers a dynamic approach to surfacing underlying stereotypical bias in large language models. Because such subtle harms are best recognized by those with lived experience, involving targets of stereotyping as red-teamers is essential. However, critical challenges remain in leveraging their lived experience for red-teaming while safeguarding psychological well-being. We conducted an empirical study of participatory red-teaming with 20 individuals stigmatized by stereotypes against non-prestigious college graduates in South Korea's rigid educational meritocracy. Through mixed-methods analysis, we found participants transformed experienced discrimination into strategic expertise for identifying biases, while facing psychological costs such as stress and negative reflections on group identity. Notably, red-team participation enhanced their sense of agency and empowerment through their role as guardians of the AI ecosystem. We discuss the implications for designing participatory red-teaming that prioritizes both the ethical treatment and the empowerment of stigmatized groups.
• Human-centered computing → Empirical studies in HCI; Empirical studies in collaborative and social computing.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 972b5b66-025d-4237-a911-8f05ad481b86Builds on20
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai et al.EMNLP 2022 · 239 citations
- The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving SupportMiriah Steiger, Timir J. Bharucha, Sukrit Venkatagiri, Martin J. Riedl et al.CHI 2021 · 168 citations
- Yes: Affirmative Consent as a Theoretical Framework for Understanding and Imagining Social PlatformsJane Im, Jill Dimond, Melody Berton, Una Lee et al.CHI 2021 · 100 citations
- AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural NuancesDhruv Agarwal, Mor Naaman, Aditya VashisthaCHI 2025 · 93 citations
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 74 citations
Related papers
- Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to EvaluationAdriana Alvarado Garcia, Ruyuan Wan, Ozioma Collins Oguine, Karla Badillo-UrquiolaCHI 2026 · 1 citation
- Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices For Generative AIBixuan Ren, Eunjeong Cheon, Jianghui LiCSCW 2025 · 3 citations
- STAR: SocioTechnical Approach to Red Teaming Language ModelsLaura Weidinger, John Mellor, Bernat Guillen Pegueroles, Nahema Marchal et al.EMNLP 2024 · 11 citations
- Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn InteractionJinchuan Zhang, Yan Zhou, Yaxin Liu, Ziming Li et al.EMNLP 2024 · 2 citations
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 13 citations
