SafetyKit: First Aid for Measuring Safety in Open-domain Conversational Systems
Emily Dinan, Gavin Abercrombie, A. Stevie Bergman, Shannon L. Spruit, Dirk Hovy, Y-Lan Boureau, Verena Rieser
Abstract
Warning: this paper contains examples that may be offensive or upsetting. The social impact of natural language processing and its applications has received increasing attention. In this position paper, we focus on the problem of safety for end-to-end conversational AI. We survey the problem landscape therein, introducing a taxonomy of three observed phenomena: the INSTIGATOR, YEA-SAYER, and IMPOSTOR effects. We then empirically assess the extent to which current tools can measure these effects and current systems display them. We release these tools as part of a "first aid kit" (SAFETYKIT) to quickly assess apparent safety concerns. Our results show that, while current tools are able to provide an estimate of the relative safety of systems in various settings, they still have several shortcomings. We suggest several future directions and discuss ethical considerations.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers8
- Gaining Wisdom from Setbacks: Aligning Large Language Models via Mistake AnalysisKai Chen, Chunwei Wang, Kuo Yang, Jianhua Han et al.ICLR 2024 · 47 citations
- ProsocialDialog: A Prosocial Backbone for Conversational AgentsHyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu et al.EMNLP 2022 · 46 citations
- Mirages. On Anthropomorphism in Dialogue SystemsGavin Abercrombie, Amanda Cercas Curry, Tanvi Dinkar, Verena Rieser et al.EMNLP 2023 · 44 citations
- SafeText: A Benchmark for Exploring Physical Safety in Language ModelsSharon Levy, Emily Allaway, Melanie Subbiah, Lydia B. Chilton et al.EMNLP 2022 · 14 citations
- Don't Forget Your ABC's: Evaluating the State-of-the-Art in Chat-Oriented Dialogue SystemsSarah E. Finch, James D. Finch, Jinho D. ChoiACL 2023 · 10 citations
Builds on10
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
- Process for Adapting Language Models to Society (PALMS) with Values-Targeted DatasetsIrene Solaiman, Christy DennisonNeurIPS 2021 · 276 citations
- MedDialog: Large-scale Medical Dialogue DatasetsGuangtao Zeng, Wenmian Yang, Zeqian Ju, Yue Yang et al.EMNLP 2020 · 163 citations
- A Distributional Approach to Controlled Text GenerationMuhammad Khalifa, Hady Elsahar, Marc DymetmanICLR 2021 · 135 citations
- Predictive Biases in Natural Language Processing Models: A Conceptual Framework and OverviewDeven Shah, H. Andrew Schwartz, Dirk HovyACL 2020 · 93 citations
Related papers
- Persona-Grounded Safety Evaluation of AI Companions in Multi-Turn ConversationsPrerna Juneja, Lika LomidzeACL 2026
- Amplifying Trans and Nonbinary Voices: A Community-Centred Harm Taxonomy for LLMsEddie L. Ungless, Sunipa Dev, Cynthia L. Bennett, Rebecca Gulotta et al.ACL 2025 · 3 citations
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu et al.FSE 2023 · 50 citations
- Is Your Multimodal Language Model Oversensitive to Safe Queries?Xirui Li, Hengguang Zhou, Ruochen Wang, Tianyi Zhou et al.ICLR 2025
- The Dark Side of AI Companionship: A Taxonomy of Harmful Algorithmic Behaviors in Human-AI RelationshipsRenwen Zhang, Han Li, Han Meng, Jinyuan Zhan et al.CHI 2025 · 122 citations
