Text Detoxification: Data Efficiency, Semantic Preservation and Model Generalization
Jing Yu, Yibo Zhao, Jiapeng Zhu, Wenming Shao, Bo Pang, Zhao Zhang, Xiang Li
摘要
The widespread dissemination of toxic content on social media poses a serious threat to both online environments and public discourse, highlighting the urgent need for detoxification methods that effectively remove toxicity while preserving the original semantics. However, existing approaches often struggle to simultaneously achieve strong detoxification performance, semantic preservation, and robustness to out-of-distribution data. Moreover, they typically rely on costly, manually annotated parallel corpora while showing poor data efficiency. To address these challenges, we propose GEM, a two-stage training framework that jointly optimizes Model Generalization, Data Efficiency, and Semantic Preservation. We first perform supervised fine-tuning on a small set of high-quality, filtered parallel data to establish a strong initialization. Then, we leverage unlabeled toxic inputs and a custom-designed reward model to train the LLM using Group Relative Policy Optimization. Experimental results demonstrate that our method effectively mitigates the trade-offs faced by previous work, achieving state-of-the-art performance with improved generalization and significantly reduced dependence on annotated data. Our code is available at https://github.com/ allacnobug/Detoxification-of-Text .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper8
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 等ACL 2022 · 被引用 96 次
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 被引用 74 次
- Text Detoxification using Large Pre-trained Neural ModelsDavid Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva 等EMNLP 2021 · 被引用 16 次
- DAC: Quantized Optimal Transport Reward-based Reinforcement Learning Approach to Detoxify Query Auto-CompletionAishwarya Maheswaran, Kaushal Kumar Maurya, Manish Gupta, Maunendra Sankar DesarkarSIGIR 2024 · 被引用 4 次
- DetoxLLM: A Framework for Detoxification with ExplanationsMd. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. LakshmananEMNLP 2024 · 被引用 4 次
相关 Paper
- Detoxification for LLM: From Dataset ItselfWei Shao, Yihang Wang, Gao yu Zhu, Ziqiang Cheng 等ACL 2026
- Unified Detoxifying and Debiasing in Language Generation via Inference-time Adaptive OptimizationZonghan Yang, Xiaoyuan Yi, Peng Li, Yang Liu 等ICLR 2023 · 被引用 6 次
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
- UniDetox: Universal Detoxification of Large Language Models via Dataset DistillationHuimin Lu, Masaru Isonuma, Junichiro Mori, Ichiro SakataICLR 2025
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 被引用 30 次
