Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
Jing Yu, Yibo Zhao, Jiapeng Zhu, Wenming Shao, Bo Pang, Zhao Zhang, Xiang Li
Abstract
The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving the original semantics. However, existing approaches often struggle to simultaneously achieve strong detoxification performance, semantic preservation, and robustness to out-of-distribution data. Moreover, they typically rely on costly, manually annotated parallel corpora while showing poor data efficiency. To address these challenges, we propose GEM, a two-stage training framework that jointly optimizes Model Generalization, Data Efficiency, and Semantic Preservation. We first perform supervised fine-tuning on a small set of high-quality, filtered parallel data to establish a strong initialization. Then, we leverage unlabeled toxic inputs and a custom-designed reward model to train the LLM using Group Relative Policy Optimization. Experimental results demonstrate that our method effectively mitigates the trade-offs faced by previous work, achieving state-of-the-art performance with improved generalization and significantly reduced dependence on annotated data. Our code is available at https://github.com/ allacnobug/Detoxification-of-Text .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers1
Ask how each one uses itBuilds on8
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy et al.ACL 2022 · 96 citations
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 74 citations
- Text Detoxification using Large Pre-trained Neural ModelsDavid Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva et al.EMNLP 2021 · 16 citations
- DAC: Quantized Optimal Transport Reward-based Reinforcement Learning Approach to Detoxify Query Auto-CompletionAishwarya Maheswaran, Kaushal Kumar Maurya, Manish Gupta, Maunendra Sankar DesarkarSIGIR 2024 · 4 citations
- DetoxLLM: A Framework for Detoxification with ExplanationsMd. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. LakshmananEMNLP 2024 · 4 citations
Related papers
- Detoxification for LLM: From Dataset ItselfWei Shao, Yihang Wang, Gao yu Zhu, Ziqiang Cheng et al.ACL 2026
- Unified Detoxifying and Debiasing in Language Generation via Inference-time Adaptive OptimizationZonghan Yang, Xiaoyuan Yi, Peng Li, Yang Liu et al.ICLR 2023 · 6 citations
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding et al.EMNLP 2024 · 1 citation
- UniDetox: Universal Detoxification of Large Language Models via Dataset DistillationHuimin Lu, Masaru Isonuma, Junichiro Mori, Ichiro SakataICLR 2025
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 30 citations
