Dark and Bright Side of Participatory Red-Teaming with Targets of Stereotyping for Eliciting Harmful Behaviors from Large Language Models
Sieun Kim, Yeeun Jo, Sungmin Na, Hyunseung Lim, Eunchae Lee, Yu Min Choi, Soohyun Cho, Hwajung Hong
摘要
Warning: This article contains stereotypical and offensive content.
Red-teaming-where adversarial prompts are crafted to expose harmful behaviors and assess risks-offers a dynamic approach to surfacing underlying stereotypical bias in large language models. Because such subtle harms are best recognized by those with lived experience, involving targets of stereotyping as red-teamers is essential. However, critical challenges remain in leveraging their lived experience for red-teaming while safeguarding psychological well-being. We conducted an empirical study of participatory red-teaming with 20 individuals stigmatized by stereotypes against non-prestigious college graduates in South Korea's rigid educational meritocracy. Through mixed-methods analysis, we found participants transformed experienced discrimination into strategic expertise for identifying biases, while facing psychological costs such as stress and negative reflections on group identity. Notably, red-team participation enhanced their sense of agency and empowerment through their role as guardians of the AI ecosystem. We discuss the implications for designing participatory red-teaming that prioritizes both the ethical treatment and the empowerment of stigmatized groups.
• Human-centered computing → Empirical studies in HCI; Empirical studies in collaborative and social computing.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper20
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- The Psychological Well-Being of Content Moderators: The Emotional Labor of Commercial Moderation and Avenues for Improving SupportMiriah Steiger, Timir J. Bharucha, Sukrit Venkatagiri, Martin J. Riedl 等CHI 2021 · 被引用 168 次
- Yes: Affirmative Consent as a Theoretical Framework for Understanding and Imagining Social PlatformsJane Im, Jill Dimond, Melody Berton, Una Lee 等CHI 2021 · 被引用 100 次
- AI Suggestions Homogenize Writing Toward Western Styles and Diminish Cultural NuancesDhruv Agarwal, Mor Naaman, Aditya VashisthaCHI 2025 · 被引用 93 次
- Is Your Toxicity My Toxicity? Exploring the Impact of Rater Identity on Toxicity AnnotationNitesh Goyal, Ian D. Kivlichan, Rachel Rosen, Lucy VassermanCSCW 2022 · 被引用 74 次
相关 Paper
- Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to EvaluationAdriana Alvarado Garcia, Ruyuan Wan, Ozioma Collins Oguine, Karla Badillo-UrquiolaCHI 2026 · 被引用 1 次
- Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices For Generative AIBixuan Ren, Eunjeong Cheon, Jianghui LiCSCW 2025 · 被引用 3 次
- STAR: SocioTechnical Approach to Red Teaming Language ModelsLaura Weidinger, John Mellor, Bernat Guillen Pegueroles, Nahema Marchal 等EMNLP 2024 · 被引用 11 次
- Holistic Automated Red Teaming for Large Language Models through Top-Down Test Case Generation and Multi-turn InteractionJinchuan Zhang, Yan Zhou, Yaxin Liu, Ziming Li 等EMNLP 2024 · 被引用 2 次
- CoP: Agentic Red-teaming for Large Language Models using Composition of PrinciplesChen Xiong, Pin-Yu Chen, Tsung-Yi HoNeurIPS 2025 · 被引用 13 次
