ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking Perturbations
Yunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei Lee
摘要
Detecting hate speech and offensive language is essential for maintaining a safe and respectful digital environment. This study examines the limitations of state-of-the-art large language models (LLMs) in identifying offensive content within systematically perturbed data, with a focus on Chinese, a language particularly susceptible to such perturbations. We introduce ToxiCloakCN 1 , an enhanced dataset derived from ToxiCN, augmented with homophonic substitutions and emoji transformations, to test the robustness of LLMs against these cloaking perturbations. Our findings reveal that existing models significantly underperform in detecting offensive content when these perturbations are applied. We provide an in-depth analysis of how different types of offensive content are affected by these perturbations and explore the alignment between human and model explanations of offensiveness. Our work highlights the urgent need for more advanced techniques in offensive language detection to combat the evolving tactics used to evade detection mechanisms.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- HalluCitation Matters: Revealing the Impact of Hallucinated References with 300 Hallucinated Papers in ACL ConferencesYusuke Sakai, Hidetaka Kamigaito, Taro WatanabeACL 2026 · 被引用 18 次
- MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking PerturbationsQiyao Xue, Yuchen Dou, Zheyuan Ryan Shi, Xiang Lorraine Li 等AAAI 2026 · 被引用 2 次
- Understanding Fanchuan in Livestreaming Platforms: A New Form of Online Antisocial BehaviorYiluo Wei, Jiahui He, Gareth TysonCSCW 2025 · 被引用 1 次
- Are Large Language Models Chronically Online Surfers? A Dataset for Chinese Internet Meme ExplanationYubo Xie, Chenkai Wang, Zongyang Ma, Fahui MiaoEMNLP 2025
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng 等EMNLP 2025
它引用的顶会 Paper5
- LoRA: Low-Rank Adaptation of Large Language ModelsEdward J. Hu, Yelong Shen, Phillip Wallis, Zeyuan Allen-Zhu 等ICLR 2022 · 被引用 18,833 次
- LLM-Adapters: An Adapter Family for Parameter-Efficient Fine-Tuning of Large Language ModelsZhiqiang Hu, Lei Wang, Yihuai Lan, Wanyu Xu 等EMNLP 2023 · 被引用 200 次
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 等EMNLP 2022 · 被引用 82 次
- RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingHui Su, Weiwei Shi, Xiaoyu Shen, Xiao Zhou 等ACL 2022 · 被引用 38 次
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min 等ACL 2023 · 被引用 25 次
相关 Paper
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding 等EMNLP 2025 · 被引用 1 次
- When Smiley Turns Hostile: Interpreting How Emojis Trigger LLMs' ToxicityShiyao Cui, Xijia Feng, Yingkang Wang, Junxiao Yang 等AAAI 2026
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes 等USENIX Security 2025
- Breaking Free from Ivory Tower: Evaluating and Enhancing Real-world Chinese Underground Adversarial Jargon DetectionZhifan Jiang, Mingxuan Liu, Yue Qin, Baojun LiuS&P 2026 · 被引用 2 次
- On the Robustness of Offensive Language ClassifiersJonathan Rusert, Zubair Shafiq, Padmini SrinivasanACL 2022 · 被引用 14 次
