Detoxifying Large Language Models via the Diversity of Toxic Samples
Ying Zhao, Yuanzhao Guo, Xuemeng Weng, Yuan Tian, Wei Wang, Yi Chang
摘要
Warning: This work contains content that may be offensive or upsetting. Eliminating toxicity from large language models (LLMs) is critical to ensure user safety. However, current methods suffer limitations in the analysis and utilization of toxic samples, failing to fully harness their potential. Through comparative analysis of toxic and safe samples, we identified that (i) toxic samples exhibit diversity and (ii) there lies specificity within this diversity. These findings suggest that leveraging these characteristics of toxic samples could enhance the performance of algorithms in LLMs detoxification. Thus, we propose a novel diverse detoxification framework, Di-vDetox, which comprises two innovative components: a Multi-Category-Induced Personalized Sample Generation (MPSG) strategy and a Scaled Contrastive Direct Preference Optimization (SC-DPO) approach. The former is designed to elicit a variety of personalized toxic responses from LLMs, while the latter is constructed to precisely and fully utilize these toxic responses. Experiments on benchmark datasets across different model scales and various detoxification tasks confirm the effectiveness of our architecture. Our codes are available at https: //github.com/zy1998-c/DivDetox .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper15
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Safe RLHF: Safe Reinforcement Learning from Human FeedbackJosef Dai, Xuehai Pan, Ruiyang Sun, Jiaming Ji 等ICLR 2024 · 被引用 656 次
相关 Paper
- From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkYuhu Shang, Xiang Cheng, Yimeng Ren, Huijia Wu 等AAAI 2026
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
- Large Language Models can Become Strong Self-DetoxifiersChing-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh 等ICLR 2025
- UniDetox: Universal Detoxification of Large Language Models via Dataset DistillationHuimin Lu, Masaru Isonuma, Junichiro Mori, Ichiro SakataICLR 2025
- Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language ModelsTassilo Klein, Moin NabiACL 2025
