Red Teaming LLMs as Socio-Technical Practice: From Exploration and Data Creation to Evaluation
Adriana Alvarado Garcia, Ruyuan Wan, Ozioma Collins Oguine, Karla Badillo-Urquiola
摘要
Recently, red teaming, with roots in security, has become a key evaluative approach to ensure the safety and reliability of Generative Artificial Intelligence. However, most existing work emphasizes technical benchmarks and attack success rates, leaving the socio-technical practices of how red teaming datasets are defined, created, and evaluated under-examined. Drawing on 22 interviews with practitioners who design and evaluate red teaming datasets, we examine the data practices and standards that underpin this work. Because adversarial datasets determine the scope and accuracy of model evaluations, they are critical artifacts for assessing potential harms from large language models. Our contributions are first, empirical evidence of practitioners conceptualizing red teaming and developing and evaluating red teaming datasets. Second, we reflect on how practitioners' conceptualization of risk leads to overlooking the context, interaction type, and user specificity. We conclude with three opportunities for HCI researchers to expand the conceptualization and data practices for red-teaming.
• Human-centered computing → Empirical studies in HCI; • Computing methodologies → Natural language generation; • Security and privacy → Human and societal aspects of security and privacy.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper14
- HarmBench: A Standardized Evaluation Framework for Automated Red Teaming and Robust RefusalMantas Mazeika, Long Phan, Xuwang Yin, Andy Zou 等ICML 2024 · 被引用 1,031 次
- "Everyone wants to do the model work, not the data work": Data Cascades in High-Stakes AINithya Sambasivan, Shivani Kapania, Hannah Highfill, Diana Akrong 等CHI 2021 · 被引用 725 次
- Red Teaming Language Models with Language ModelsEthan Perez, Saffron Huang, H. Francis Song, Trevor Cai 等EMNLP 2022 · 被引用 239 次
- Do Datasets Have Politics? Disciplinary Values in Computer Vision Dataset DevelopmentMorgan Klaus Scheuerman, Alex Hanna, Emily DentonCSCW 2021 · 被引用 169 次
- Between Subjectivity and Imposition: Power Dynamics in Data Annotation for Computer VisionMilagros Miceli, Martin Schuessler, Tianling YangCSCW 2020 · 被引用 148 次
相关 Paper
- Organization Matters: A Qualitative Study of Organizational Dynamics in Red Teaming Practices For Generative AIBixuan Ren, Eunjeong Cheon, Jianghui LiCSCW 2025 · 被引用 3 次
- Emerging Data Practices: Data Work in the Era of Large Language ModelsAdriana Alvarado Garcia, Heloisa Candello, Karla Badillo-Urquiola, Marisol Wong-VillacresCHI 2025 · 被引用 6 次
- StealthGraph: Exposing Domain-Specific Risks in LLMs through Knowledge-Graph-Guided Harmful Prompt GenerationHuawei Zheng, Xinqi Jiang, Sen Yang, Shouling Ji 等ACL 2026 · 被引用 1 次
- A Coin Flip for Safety: LLM Judges Fail to Reliably Measure Adversarial RobustnessLeo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami 等ICML 2026 · 被引用 15 次
- Automated Red Teaming with GOAT: the Generative Offensive Agent TesterMaya Pavlova, Erik Brinkman, Krithika Iyer, Vítor Albiero 等ICML 2025
