UniDetox: Universal Detoxification of Large Language Models via Dataset Distillation
Huimin Lu, Masaru Isonuma, Junichiro Mori, Ichiro Sakata
摘要
We present UNIDETOX, a universally applicable method designed to mitigate toxicity across various large language models (LLMs). Previous detoxification methods are typically model-specific, addressing only individual models or model families, and require careful hyperparameter tuning due to the trade-off between detoxification efficacy and language modeling performance. In contrast, UNIDETOX provides a detoxification technique that can be universally applied to a wide range of LLMs without the need for separate model-specific tuning. Specifically, we propose a novel and efficient dataset distillation technique for detoxification using contrastive decoding. This approach distills detoxifying representations in the form of synthetic text data, enabling universal detoxification of any LLM through fine-tuning with the distilled text. Our experiments demonstrate that the detoxifying text distilled from GPT-2 can effectively detoxify larger models, including OPT, Falcon, and LLaMA-2. Furthermore, UNIDETOX eliminates the need for separate hyperparameter tuning for each model, as a single hyperparameter configuration can be seamlessly applied across different models. Additionally, analysis of the detoxifying text reveals a reduction in politically biased content, providing insights into the attributes necessary for effective detoxification of LLMs. Our codes are available at https://github.com/EminLU/UniDetox .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Beyond Modality Collapse: Representation Blending for Multimodal Dataset DistillationXin Zhang, Ziruo Zhang, Jiawei Du, Zuozhu Liu 等NeurIPS 2025 · 被引用 9 次
- Test-Time Detoxification without Training or Learning AnythingBaturay Saglam, Dionysios KalogeriasICML 2026 · 被引用 2 次
- Generative Data Transformation: From Mixed to Unified DataJiaqing Zhang, Mingjia Yin, Hao Wang, Yuxin Tian 等WWW 2026
- From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkYuhu Shang, Xiang Cheng, Yimeng Ren, Huijia Wu 等AAAI 2026
- Detoxification for LLM: From Dataset ItselfWei Shao, Yihang Wang, Gao yu Zhu, Ziqiang Cheng 等ACL 2026
它引用的顶会 Paper27
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- Measuring Massive Multitask Language UnderstandingDan Hendrycks, Collin Burns, Steven Basart, Andy Zou 等ICLR 2021 · 被引用 7,905 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
相关 Paper
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 被引用 30 次
- Large Language Models can Become Strong Self-DetoxifiersChing-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh 等ICLR 2025
- DetoxLLM: A Framework for Detoxification with ExplanationsMd. Tawkat Islam Khondaker, Muhammad Abdul-Mageed, Laks V. S. LakshmananEMNLP 2024 · 被引用 4 次
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
- Text Detoxification: Data Efficiency, Semantic Preservation and Model GeneralizationJing Yu, Yibo Zhao, Jiapeng Zhu, Wenming Shao 等EMNLP 2025
