Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive Contexts
Ashutosh Baheti, Maarten Sap, Alan Ritter, Mark O. Riedl
摘要
Dialogue models trained on human conversations inadvertently learn to generate toxic responses. In addition to producing explicitly offensive utterances, these models can also implicitly insult a group or individual by aligning themselves with an offensive statement. To better understand the dynamics of contextually offensive language, we investigate the stance of dialogue model responses in offensive Reddit conversations. Specifically, we create TOXICHAT, a crowd-annotated dataset of 2,000 Reddit threads and model responses labeled with offensive language and stance. Our analysis reveals that 42% of human responses agree with toxic comments, whereas only 13% agree with safe comments. This undesirable behavior is learned by neural dialogue models, such as DialoGPT, which we show are two times more likely to agree with offensive comments. To enable automatic detection of offensive language, we fine-tuned transformerbased classifiers on TOXICHAT that achieve 0.71 F 1 for offensive labels and 0.53 Macro-F 1 for stance labels. Finally, we quantify the effectiveness of controllable text generation (CTG) methods to mitigate the tendency of neural dialogue models to agree with offensive comments. Compared to the baseline, our best CTG model achieves a 19% reduction in agreement with offensive comments and produces 29% fewer offensive replies. Our work highlights the need for further efforts to characterize and analyze inappropriate behavior in dialogue models, in order to help make them safer. 1
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper22
- Mix and Match: Learning-free Controllable Text Generationusing Energy Language ModelsFatemehsadat Mireshghallah, Kartik Goyal, Taylor Berg-KirkpatrickACL 2022 · 被引用 90 次
- Exploring the Limits of Domain-Adaptive Training for Detoxifying Large-Scale Language ModelsBoxin Wang, Wei Ping, Chaowei Xiao, Peng Xu 等NeurIPS 2022 · 被引用 89 次
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 等EMNLP 2022 · 被引用 82 次
- SODA: Million-scale Dialogue Distillation with Social Commonsense ContextualizationHyunwoo Kim, Jack Hessel, Liwei Jiang, Peter West 等EMNLP 2023 · 被引用 60 次
- BiasAsker: Measuring the Bias in Conversational AI SystemYuxuan Wan, Wenxuan Wang, Pinjia He, Jiazhen Gu 等FSE 2023 · 被引用 50 次
它引用的顶会 Paper11
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Don't Stop Pretraining: Adapt Language Models to Domains and TasksSuchin Gururangan, Ana Marasovic, Swabha Swayamdipta, Kyle Lo 等ACL 2020 · 被引用 93 次
- Hierarchical Reinforcement Learning for Open-Domain DialogAbdelrhman Saleh, Natasha Jaques, Asma Ghandeharioun, Judy Hanwen Shen 等AAAI 2020 · 被引用 60 次
- Mitigating Gender Bias for Neural Dialogue Generation with Adversarial LearningHaochen Liu, Wentao Wang, Yiqi Wang, Hui Liu 等EMNLP 2020 · 被引用 55 次
- Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media ConversationsJianfei Yu, Jing Jiang, Ling Min Serena Khoo, Hai Leong Chieu 等EMNLP 2020 · 被引用 49 次
相关 Paper
- Why So Toxic?: Measuring and Triggering Toxic Behavior in Open-Domain ChatbotsWai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro 等CCS 2022 · 被引用 34 次
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi 等ACL 2023 · 被引用 5 次
- Language Detoxification with Attribute-Discriminative Latent SpaceJin Myung Kwak, Minseon Kim, Sung Ju HwangACL 2023 · 被引用 2 次
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding 等EMNLP 2025 · 被引用 1 次
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap 等ACL 2022
