Text Detoxification using Large Pre-trained Neural Models
David Dale, Anton Voronov, Daryna Dementieva, Varvara Logacheva, Olga Kozlova, Nikita Semenov, Alexander Panchenko
摘要
We present two novel unsupervised methods for eliminating toxicity in text. Our first method combines two recent ideas: (1) guidance of the generation process with small styleconditional language models and (2) use of paraphrasing models to perform style transfer. We use a well-performing paraphraser guided by style-trained language models to keep the text content and remove toxicity. Our second method uses BERT to replace toxic words with their non-offensive synonyms. We make the method more flexible by enabling BERT to replace mask tokens with a variable number of words. Finally, we present the first largescale comparative study of style transfer models on the task of toxicity removal. We compare our models with a number of methods for style transfer. The models are evaluated in a reference-free way using a combination of unsupervised style transfer metrics. Both methods we suggest yield new SOTA results.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper19
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 等ACL 2022 · 被引用 96 次
- You Only Prompt Once: On the Capabilities of Prompt Learning on Large Language Models to Tackle Toxic ContentXinlei He, Savvas Zannettou, Yun Shen, Yang ZhangS&P 2024 · 被引用 74 次
- ParaGuide: Guided Diffusion Paraphrasers for Plug-and-Play Textual Style TransferZachary Horvitz, Ajay Patel, Chris Callison-Burch, Zhou Yu 等AAAI 2024 · 被引用 21 次
- MERA: A Comprehensive LLM Evaluation in RussianAlena Fenogenova, Artem Chervyakov, Nikita Martynov, Anastasia Kozlova 等ACL 2024 · 被引用 10 次
- IF-Guide: Influence Function-Guided Detoxification of LLMsZachary Coalson, Juhan Bae, Nicholas Carlini, Sanghyun HongNeurIPS 2025 · 被引用 9 次
它引用的顶会 Paper5
- Plug and Play Language Models: A Simple Approach to Controlled Text GenerationSumanth Dathathri, Andrea Madotto, Janice Lan, Jane Hung 等ICLR 2020 · 被引用 1,166 次
- A Probabilistic Formulation of Unsupervised Text Style TransferJunxian He, Xinyi Wang, Graham Neubig, Taylor Berg-KirkpatrickICLR 2020 · 被引用 136 次
- Style-transfer and Paraphrase: Looking for a Sensible Semantic Similarity MetricIvan P. Yamshchikov, Viacheslav Shibaev, Nikolay Khlebnikov, Alexey TikhonovAAAI 2021 · 被引用 38 次
- Reformulating Unsupervised Style Transfer as Paraphrase GenerationKalpesh Krishna, John Wieting, Mohit IyyerEMNLP 2020 · 被引用 9 次
- Politeness Transfer: A Tag and Generate ApproachAman Madaan, Amrith Setlur, Tanmay Parekh, Barnabás Póczos 等ACL 2020 · 被引用 6 次
相关 Paper
- Masked Based Unsupervised Content TransferRon Mokady, Sagie Benaim, Lior Wolf, Amit BermanoICLR 2020
- Generic resources are what you need: Style transfer tasks without task-specific parallel training dataHuiyuan Lai, Antonio Toral, Malvina NissimEMNLP 2021 · 被引用 14 次
- Leashing the Inner Demons: Self-Detoxification for Language ModelsCanwen Xu, Zexue He, Zhankui He, Julian J. McAuleyAAAI 2022 · 被引用 30 次
- Prompt-and-Rerank: A Method for Zero-Shot and Few-Shot Arbitrary Textual Style Transfer with Small Language ModelsMirac Suzgun, Luke Melas-Kyriazi, Dan JurafskyEMNLP 2022 · 被引用 34 次
- Lexical Simplification with Pretrained EncodersJipeng Qiang, Yun Li, Yi Zhu, Yunhao Yuan 等AAAI 2020 · 被引用 86 次
