Contrastive Perplexity for Controlled Generation: An Application in Detoxifying Large Language Models
Tassilo Klein, Moin Nabi
摘要
The generation of toxic content by large language models (LLMs) remains a critical challenge for the safe deployment of language technology. We propose a novel framework for implicit knowledge editing and controlled text generation by fine-tuning LLMs with a prototype-based contrastive perplexity objective. Central to our method is the construction of hard negatives-toxic outputs that are generated through adversarial paraphrasing to be semantically similar and model probability to their non-toxic counterparts. By training on these challenging and realistic pairs, our approach ensures robust and stable contrastive optimization. Experimental results in the domain of detoxification demonstrate that our method significantly reduces toxic generation while maintaining strong performance on downstream tasks such as commonsense reasoning and reading comprehension. Our findings highlight the effectiveness of exploiting hard negatives for attribute-aware fine-tuning. 1 Disclaimer: Contains sensitive content.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper1
问问它们各自怎么用它它引用的顶会 Paper22
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- Direct Preference Optimization: Your Language Model is Secretly a Reward ModelRafael Rafailov, Archit Sharma, Eric Mitchell, Christopher D. Manning 等NeurIPS 2023 · 被引用 10,924 次
- The Curious Case of Neural Text DegenerationAri Holtzman, Jan Buys, Li Du, Maxwell Forbes 等ICLR 2020 · 被引用 4,112 次
- WinoGrande: An Adversarial Winograd Schema Challenge at ScaleKeisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, Yejin ChoiAAAI 2020 · 被引用 3,037 次
- PIQA: Reasoning about Physical Commonsense in Natural LanguageYonatan Bisk, Rowan Zellers, Ronan Le Bras, Jianfeng Gao 等AAAI 2020 · 被引用 2,916 次
相关 Paper
- Self-Detoxifying Language Models via Toxification ReversalChak Tou Leong, Yi Cheng, Jiashuo Wang, Jian Wang 等EMNLP 2023 · 被引用 12 次
- From Chaos to Cure: A Prefix Heuristics Guided Model-Agnostic Adaptive Detoxification FrameworkYuhu Shang, Xiang Cheng, Yimeng Ren, Huijia Wu 等AAAI 2026
- CMD: a framework for Context-aware Model self-DetoxificationZecheng Tang, Keyan Zhou, Juntao Li, Yuyang Ding 等EMNLP 2024 · 被引用 1 次
- Detoxifying Large Language Models via Autoregressive Reward Guided Representation EditingYisong Xiao, Aishan Liu, Siyuan Liang, Zonghao Ying 等NeurIPS 2025 · 被引用 12 次
- Language Detoxification with Attribute-Discriminative Latent SpaceJin Myung Kwak, Minseon Kim, Sung Ju HwangACL 2023 · 被引用 2 次
