Enhancing Chinese Offensive Language Detection with Homophonic Perturbation
Junqi Wu, Shujie Ji, Kang Zhong, Huiling Peng, Zhendongxiao, Xiongding Liu, Wu Wei
Abstract
Detecting offensive language in Chinese is challenging due to homophonic substitutions used to evade detection. We propose a framework to improve large language models'robustness against such phonetic attacks. First, we construct HED-COLD 1 , the first largescale and systematic homophonic dataset for Chinese offensive language detection. Additionally, we design a homophone-aware pretraining strategy that learns the mappings among orthography, phonetics, and semantics between original and perturbed text. Experimental results show that our approach achieves state-of-the-art performance on both the COLD test set and the toxicity benchmark ToxiCloakCN. Notably, it achieves greater gains in domains susceptible to homophonic attacks, such as gender and regional content. These results demonstrate improved robustness and generalization against phonetic adversarial attacks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 40c31487-41b9-4cf6-8315-ff6426b6b6c4Builds on5
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng et al.EMNLP 2022 · 82 citations
- RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingHui Su, Weiwei Shi, Xiaoyu Shen, Xiao Zhou et al.ACL 2022 · 38 citations
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min et al.ACL 2023 · 25 citations
- CSCD-NS: a Chinese Spelling Check Dataset for Native SpeakersYong Hu, Fandong Meng, Jie ZhouACL 2024 · 9 citations
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 5 citations
Related papers
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding et al.EMNLP 2025 · 1 citation
- Towards Robust Detection of Chinese Toxic Variants via Dynamic Knowledge Graph-LLM ReasoningShaochen Yang, Kefei Zhou, Wei XuWWW 2026
- Breaking Free from Ivory Tower: Evaluating and Enhancing Real-world Chinese Underground Adversarial Jargon DetectionZhifan Jiang, Mingxuan Liu, Yue Qin, Baojun LiuS&P 2026 · 2 citations
- TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine TranslationJinfeng Li, Tianyu Du, Shouling Ji, Rong Zhang et al.USENIX Security 2020
- HateBench: Benchmarking Hate Speech Detectors on LLM-Generated Content and Hate CampaignsXinyue Shen, Yixin Wu, Yiting Qu, Michael Backes et al.USENIX Security 2025
