Leashing the Inner Demons: Self-Detoxification for Language Models
Canwen Xu, Zexue He, Zhankui He, Julian J. McAuley
Abstract
Language models (LMs) can reproduce (or amplify) toxic language seen during training, which poses a risk to their practical application. In this paper, we conduct extensive experiments to study this phenomenon. We analyze the impact of prompts, decoding strategies and training corpora on the output toxicity. Based on our findings, we propose a simple yet effective unsupervised method for language models to ``detoxify'' themselves without an additional large corpus or external discriminator. Compared to a supervised baseline, our proposed method shows better toxicity reduction with good generation quality in the generated content under multiple settings. Warning: some examples shown in the paper may contain uncensored offensive content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext babe7898-31ad-469b-8a24-795c55d3f9c1Cited by top-tier papers12
- Why So Toxic?: Measuring and Triggering Toxic Behavior in Open-Domain ChatbotsWai Man Si, Michael Backes, Jeremy Blackburn, Emiliano De Cristofaro et al.CCS 2022 · 34 citations
- Self-Detoxifying Language Models via Toxification ReversalChak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang et al.EMNLP 2023 · 12 citations
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 9 citations
- Interactive Text GenerationFelix Faltings, Michel Galley, Kianté Brantley, Baolin Peng et al.EMNLP 2023 · 7 citations
- MedEval: A Multi-Level, Multi-Task, and Multi-Domain Medical Benchmark for Language Model EvaluationZexue He, Yu Wang, An Yan, Yao Liu et al.EMNLP 2023 · 7 citations
Builds on8
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes et al.ICLR 2020 · 4,112 citations
- The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural NetworksNicholas Carlini, Chang Liu, Úlfar Erlingsson, Jernej Kos et al.USENIX Security 2019 · 1,386 citations
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung et al.ICLR 2020 · 1,166 citations
Related papers
- Exploring the Limits of Domain-Adaptive Training for Detoxifying Large-Scale Language ModelsBoxin Wang, Wei Ping, Chaowei Xiao, Peng Xu et al.NeurIPS 2022 · 89 citations
- MIL-Decoding: Detoxifying Language Models at Token-Level via Multiple Instance LearningXu Zhang, Xiaojun WanACL 2023 · 3 citations
- Large Language Models can Become Strong Self-DetoxifiersChing-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh et al.ICLR 2025
- Language Detoxification with Attribute-Discriminative Latent SpaceJin Myung Kwak, Minseon Kim, Sung Ju HwangACL 2023 · 2 citations
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding et al.EMNLP 2024 · 1 citation
