Chinese Toxic Language Mitigation via Sentiment Polarity Consistent Rewrites
Xintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding, Longyue Wang, Chris Biemann
摘要
Detoxifying offensive language while preserving the speaker's original intent is a challenging yet critical goal for improving the quality of online interactions. Although large language models (LLMs) show promise in rewriting toxic content, they often default to overly polite rewrites, distorting the emotional tone and communicative intent. This problem is especially acute in Chinese, where toxicity often arises implicitly through emojis, homophones, or discourse context. We present TOXIREWRITECN, the first Chinese detoxification dataset explicitly designed to preserve sentiment polarity. The dataset comprises 1,556 carefully annotated triplets, each containing a toxic sentence, a sentiment-aligned non-toxic rewrite, and labeled toxic spans. It covers five real-world scenarios: standard expressions, emoji-induced and homophonic toxicity, as well as single-turn and multi-turn dialogues. We evaluate 17 LLMs, including commercial and open-source models with variant architectures, across four dimensions: detoxification accuracy, fluency, content preservation, and sentiment polarity. Results show that while commercial and MoE models perform best overall, all models struggle to balance safety with emotional fidelity in more subtle or context-heavy settings such as emoji, homophone, and dialogue-based inputs. We release TOXIREWRITECN to support future research on controllable, sentiment-aware detoxification for Chinese. Caution: This paper contains examples of violent or offensive language that may be disturbing to some readers.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- ParaDetox: Detoxification with Parallel DataVarvara Logacheva, Daryna Dementieva, Sergey Ustyantsev, Daniil Moskovskiy 等ACL 2022 · 被引用 96 次
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 等EMNLP 2022 · 被引用 82 次
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min 等ACL 2023 · 被引用 25 次
- Toxicity Detection is NOT all you Need: Measuring the Gaps to Supporting Volunteer Content Moderators through a User-Centric MethodYang Trista Cao, Lovely-Frances Domingo, Sarah A. Gilbert, Michelle L. Mazurek 等EMNLP 2024 · 被引用 4 次
- Walking in Others' Shoes: How Perspective-Taking Guides Large Language Models in Reducing Toxicity and BiasRongwu Xu, Zi'an Zhou, Tianwei Zhang, Zehan Qi 等EMNLP 2024 · 被引用 2 次
相关 Paper
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 被引用 5 次
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng 等EMNLP 2025
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' ToxicityShiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang 等AAAI 2026
- UniDetox: Universal Detoxification of Large Language Models via Dataset DistillationHuimin Lu, Masaru Isonuma, Junichiro Mori, Ichiro SakataICLR 2025
- MIL-Decoding: Detoxifying Language Models at Token-Level via Multiple Instance LearningXu Zhang, Xiaojun WanACL 2023 · 被引用 3 次
