Leashing the Inner Demons: Self-Detoxification for Language Models
Canwen Xu, Zexue He, Zhankui He, Julian J. McAuley
摘要
Language models (LMs) can reproduce (or amplify) toxic language seen during training, which poses a risk to their practical application. In this paper, we conduct extensive experiments to study this phenomenon. We analyze the impact of prompts, decoding strategies and training corpora on the output toxicity. Based on our findings, we propose a simple yet effective unsupervised method for language models to ``detoxify'' themselves without an additional large corpus or external discriminator. Compared to a supervised baseline, our proposed method shows better toxicity reduction with good generation quality in the generated content under multiple settings. Warning: some examples shown in the paper may contain uncensored offensive content.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper12
- Why So Toxic?: Measuring and Triggering Toxic Behavior in Open-Domain ChatbotsWai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro 等CCS 2022 · 被引用 34 次
- Self-Detoxifying Language Models via Toxification ReversalChak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang 等EMNLP 2023 · 被引用 12 次
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 被引用 9 次
- Interactive Text GenerationFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng 等EMNLP 2023 · 被引用 7 次
- MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationZexue He, Yu Wang, An Yan, Yao Liu 等EMNLP 2023 · 被引用 7 次
它引用的顶会 Paper8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos 等USENIX Security 2019 · 被引用 1,386 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- Exploring the Limits of Domain-Adaptive Training for Detoxifying Large-Scale Language ModelsBoxin Wang, Wei Ping, Chaowei Xiao, Peng Xu 等NeurIPS 2022 · 被引用 89 次
- MIL-Decoding: Detoxifying Language Models at Token-Level via Multiple Instance LearningXu Zhang, Xiaojun WanACL 2023 · 被引用 3 次
- Large Language Models can Become Strong Self-DetoxifiersChing-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh 等ICLR 2025
- Language Detoxification with Attribute-Discriminative Latent SpaceJin Myung Kwak, Minseon Kim, Sung Ju HwangACL 2023 · 被引用 2 次
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
