Lune

EMNLP2025Top-tier venue

Detoxifying Large Language Models via the Diversity of Toxic Samples

Ying Zhao, Yuanzhao Guo, Xuemeng Weng, Yuan Tian, Wei Wang, Yi Chang

2025Year
1Top-tier citations

Abstract

Warning: This work contains content that may be offensive or upsetting. Eliminating toxicity from large language models (LLMs) is critical to ensure user safety. However, current methods suffer limitations in the analysis and utilization of toxic samples, failing to fully harness their potential. Through comparative analysis of toxic and safe samples, we identified that (i) toxic samples exhibit diversity and (ii) there lies specificity within this diversity. These findings suggest that leveraging these characteristics of toxic samples could enhance the performance of algorithms in LLMs detoxification. Thus, we propose a novel diverse detoxification framework, Di-vDetox, which comprises two innovative components: a Multi-Category-Induced Personalized Sample Generation (MPSG) strategy and a Scaled Contrastive Direct Preference Optimization (SC-DPO) approach. The former is designed to elicit a variety of personalized toxic responses from LLMs, while the latter is constructed to precisely and fully utilize these toxic responses. Experiments on benchmark datasets across different model scales and various detoxification tasks confirm the effectiveness of our architecture. Our codes are available at https: //github.com/zy1998-c/DivDetox .

Ask about this paper

Your agent reads all of it.

Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.

Questions to start from

Your agent calls

Luneget_paper_fulltext

Ask in Lune

Free to start. No credit card required.

lune papers fulltext e5a00bdf-fece-4503-8bb8-c359074ed720

Cited by top-tier papers1

Ask how each one uses it

Builds on15

Related papers

Dusk over the sea between two cliffs drawn in fine vertical lines