Self-Detoxifying Language Models via Toxification Reversal
Chak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang, Wenjie Li
摘要
Language model detoxification aims to minimize the risk of generating offensive or harmful content in pretrained language models (PLMs) for safer deployment. Existing methods can be roughly categorized as finetuning-based and decoding-based. However, the former is often resource-intensive, while the latter relies on additional components and potentially compromises the generation fluency. In this paper, we propose a more lightweight approach that enables the PLM itself to achieve "selfdetoxification". Our method is built upon the observation that prepending a negative steering prompt can effectively induce PLMs to generate toxic content. At the same time, we are inspired by the recent research in the interpretability field, which formulates the evolving contextualized representations within the PLM as an information stream facilitated by the attention layers. Drawing on this idea, we devise a method to identify the toxification direction from the normal generation process to the one prompted with the negative prefix, and then steer the generation to the reversed direction by manipulating the information movement within the attention layers. Experimental results show that our approach, without any fine-tuning or extra components, can achieve comparable performance with state-of-the-art methods. 1 1 Code is available at https://github.com/ cooperleong00/ToxificationReversal *Equal contribution Toxic Non-Toxic Context Embeddings Toxification Process Generation Space MHSA MHSA MHSA MHSA MHSA MHSA Toxification Direction Toxification Reversal Direction
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper26
- In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space SteeringSheng Liu, Haotian Ye, Lei Xing, James Y. ZouICML 2024 · 被引用 244 次
- Fundamental Limitations of Alignment in Large Language ModelsYotam Wolf, Noam Wies, Oshri Avnery, Yoav Levine 等ICML 2024 · 被引用 186 次
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingYisong Xiao, Aishan Liu, Siyuan Liang, Zonghao Ying 等NeurIPS 2025 · 被引用 12 次
- Spherical Steering: Geometry-Aware Activation Rotation for Language ModelsZejia You, Chunyuan Deng, Hanjie ChenICML 2026 · 被引用 12 次
- Steering When Necessary: Flexible Steering Large Language Models with BacktrackingZifeng Cheng, Jinwei Gan, Zhiwei Jiang, Cong Wang 等NeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper16
- Training language models to follow instructions with human feedbackLong Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida 等NeurIPS 2022 · 被引用 24,707 次
- Finetuned Language Models are Zero-Shot LearnersJason Wei, Maarten Bosma, Vincent Y. Zhao, Kelvin Guu 等ICLR 2022 · 被引用 4,966 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- Fine-Tuning can Distort Pretrained Features and Underperform Out-of-DistributionAnanya Kumar, Aditi Raghunathan, Robbie Matthew Jones, Tengyu Ma 等ICLR 2022 · 被引用 911 次
相关 Paper
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 被引用 30 次
- DSCD: Large Language Model Detoxification with Self-Constrained DecodingMing Dong, Jinkui Zhang, Bolong Zheng, Xinhui Tu 等EMNLP 2025 · 被引用 1 次
- Large Language Models can Become Strong Self-DetoxifiersChing-Yun Ko, Pin-Yu Chen, Payel Das, Youssef Mroueh 等ICLR 2025
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
- Detoxification for LLM: From Dataset ItselfWei Shao, Yihang Wang, Gao yu Zhu, Ziqiang Cheng 等ACL 2026
