CMD: a framework for Context-aware Model self-Detoxification
Zecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding, Pinzheng Wang, Yan Bowen, Renjie Hua, Min Zhang
Abstract
Text detoxification aims to minimize the risk of language models producing toxic content. However, existing detoxification methods fail to balance the detoxification effectiveness and generation quality. This issue arises from neglecting the constraints imposed by the context: language models are designed to generate output that closely matches the given context, while detoxification methods strive to ensure the safety of the output, even if it deviates semantically from the context. Given this, we introduce a Context-aware Model self-Detoxification (CMD) framework that pays attention to both the context and the detoxification process, i.e., first detoxifying the context and then making the language model generate along the safe context. Specifically, CMD framework involves two phases: utilizing language models to synthesize data and applying these data for training. We also introduce a toxic contrastive loss that encourages the model generation away from the negative toxic samples. Experiments on various LLMs have verified the effectiveness of our MSD framework, which can yield the best performance compared to baselines. 1 Warning: cases in this paper may contain offensive content.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language ModelsTassilo Klein, Moin NabiACL 2025
- From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkYuhu Shang, Xiang Cheng, Yimeng Ren, Huijia Wu et al.AAAI 2026
- Detoxification for LLM: From Dataset ItselfWei Shao, Yihang Wang, Gao yu Zhu, Ziqiang Cheng et al.ACL 2026
Builds on18
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah et al.NeurIPS 2020 · 64,255 citations
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida et al.NeurIPS 2022 · 24,707 citations
- Chain-of-Thought Prompting Elicits Reasoning in Large Language ModelsJason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma et al.NeurIPS 2022 · 22,562 citations
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu et al.ICLR 2022 · 18,833 citations
- BERTScore: Evaluating Text Generation with BERTTianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger et al.ICLR 2020 · 8,443 citations
Related papers
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 30 citations
- DSCD: Large Language Model Detoxification with Self-Constrained DecodingMing Dong, Jinkui Zhang, Bolong Zheng, Xinhui Tu et al.EMNLP 2025 · 1 citation
- Self-Detoxifying Language Models via Toxification ReversalChak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang et al.EMNLP 2023 · 12 citations
- Detoxifying Large Language Models via the Diversity of Toxic SamplesYing Zhao, Yuanzhao Guo, Xuemeng Weng, Yuan Tian et al.EMNLP 2025
- DetoxLLM: A Framework for Detoxification with ExplanationsMd. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. LakshmananEMNLP 2024 · 4 citations
