COLD: A Benchmark for Chinese Offensive Language Detection
Jiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng, Fei Mi, Helen Meng, Minlie Huang
Abstract
Offensive language detection is increasingly crucial for maintaining a civilized social media platform and deploying pre-trained language models. However, this task in Chinese is still under exploration due to the scarcity of reliable datasets. To this end, we propose a benchmark -COLD for Chinese offensive language analysis, including a Chinese Offensive Language Dataset -COLDATASET and a baseline detector -COLDETECTOR which is trained on the dataset. We show that the COLD benchmark contributes to Chinese offensive language detection which is challenging for existing resources. We then deploy the COLDETECTOR and conduct detailed analyses on popular Chinese pre-trained language models. We first analyze the offensiveness of existing generative models and show that these models inevitably expose varying degrees of offensive issues. Furthermore, we investigate the factors that influence the offensive generations, and we find that anti-bias contents and keywords referring to certain groups or revealing negative attitudes trigger offensive outputs easier. 1
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 81da181d-11f8-4ada-8246-b0303c9de4e5Cited by top-tier papers19
- SafeDialBench: A Fine-Grained Safety Evaluation Benchmark for Large Language Models in Multi-Turn Dialogues with Diverse Jailbreak AttacksHongye Cao, Sijia Jing, Yanming Wang, Ziyue Peng et al.ICLR 2026 · 28 citations
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min et al.ACL 2023 · 25 citations
- MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and BilibiliHan Wang, Tan Rui Yang, Usman Naseem, Roy Ka-Wei LeeACM MM 2024 · 23 citations
- DeMod: A Holistic Tool with Explainable Detection and Personalized Modification for Toxicity CensorshipYaqiong Li, Peng Zhang, Hansu Gu, Tun Lu et al.CSCW 2025 · 6 citations
- SafeConv: Explaining and Correcting Conversational Unsafe BehaviorMian Zhang, Lifeng Jin, Linfeng Song, Haitao Mi et al.ACL 2023 · 5 citations
Builds on4
- Just Say No: Analyzing the Stance of Neural Dialogue Generation in Offensive ContextsAshutosh Baheti, Maarten Sap, Alan Ritter, Mark O. RiedlEMNLP 2021 · 50 citations
- CrowS-Pairs: A Challenge Dataset for Measuring Social Biases in Masked Language ModelsNikita Nangia, Clara Vania, Rasika Bhalerao, Samuel R. BowmanEMNLP 2020 · 19 citations
- Probing Toxic Content in Large Pre-Trained Language ModelsNedjma Ousidhoum, Xinran Zhao, Tianqing Fang, Yangqiu Song et al.ACL 2021
- StereoSet: Measuring stereotypical bias in pretrained language modelsMoin Nadeem, Anna Bethke, Siva ReddyACL 2021
Related papers
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng et al.EMNLP 2025
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 5 citations
- KOLD: Korean Offensive Language DatasetYounghoon Jeong, Juhyun Oh, Jongwon Lee, Jaimeen Ahn et al.EMNLP 2022 · 41 citations
- Agreeing to Disagree: Annotating Offensive Language Datasets with Annotators' DisagreementElisa Leonardelli, Stefano Menini, Alessio Palmero Aprosio, Marco Guerini et al.EMNLP 2021 · 2 citations
- GeWu: A Culturally-Grounded Chinese Benchmark for Multi-Stage Social Bias Evaluation in Large Language ModelsYi Lin, Ziyi Zhou, Jiashi Gao, Xinwei Guo et al.AAAI 2026
