SafeConv: Explaining and Correcting Conversational Unsafe Behavior
Mian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi, Wenliang Chen, Dong Yu
摘要
One of the main challenges open-domain end-to-end dialogue systems, or chatbots, face is the prevalence of unsafe behavior, such as toxic languages and harmful suggestions. However, existing dialogue datasets do not provide enough annotation to explain and correct such unsafe behavior. In this work, we construct a new dataset called SafeConv for the research of conversational safety: (1) Besides the utterance-level safety labels, SafeConv also provides unsafe spans in an utterance, information able to indicate which words contribute to the detected unsafe behavior; (2) SafeConv provides safe alternative responses to continue the conversation when unsafe behavior detected, guiding the conversation to a gentle trajectory. By virtue of the comprehensive annotation of SafeConv, we benchmark three powerful models for the mitigation of conversational unsafe behavior, including a checker to detect unsafe utterances, a tagger to extract unsafe spans, and a rewriter to convert an unsafe response to a safe version. Moreover, we explore the huge benefits brought by combining the models for explaining the emergence of unsafe behavior and detoxifying chatbots. Experiments show that the detected unsafe behavior could be well explained with unsafe spans and popular chatbots could be detoxified by a huge extent. The dataset is available at https://github.com/mianzhang/SafeConv.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- LongLoRA: Efficient Fine-tuning of Long-Context Large Language ModelsYukang Chen, Shengju Qian, Haotian Tang, Xin Lai 等ICLR 2024 · 被引用 254 次
- CounselBench: A Large-Scale Expert Evaluation and Adversarial Benchmarking of Large Language Models in Mental Health Question AnsweringYahan Li, Jifan Yao, John Bosco S. Bunyi, Adam C. Frank 等ICLR 2026 · 被引用 24 次
- Seeing the Image: Prioritizing Visual Correlation by Contrastive AlignmentXin Xiao, Bohong Wu, Jiacong Wang, Chunyuan Li 等NeurIPS 2024 · 被引用 24 次
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
- Nullspace Disentanglement for Red Teaming Language ModelsYi Han, Yuanxing Liu, Weinan Zhang, Ting LiuEMNLP 2025
它引用的顶会 Paper9
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and ComprehensionMike Lewis, Yinhan Liu, Naman Goyal, Marjan Ghazvininejad 等ACL 2020 · 被引用 1,224 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 等EMNLP 2022 · 被引用 82 次
相关 Paper
- ProsocialDialog: A Prosocial Backbone for Conversational AgentsHyunwoo Kim, Youngjae Yu, Liwei Jiang, Ximing Lu 等EMNLP 2022 · 被引用 46 次
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding 等EMNLP 2025 · 被引用 1 次
- SaFeRDialogues: Taking Feedback Gracefully after Conversational Safety FailuresMegan Ung, Jing Xu, Y-Lan BoureauACL 2022 · 被引用 54 次
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 等ACL 2022 · 被引用 96 次
- ConvAbuse: Data, Analysis, and Benchmarks for Nuanced Detection in Conversational AIAmanda Cercas Curry, Gavin Abercrombie, Verena RieserEMNLP 2021 · 被引用 38 次
