Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and Benchmarks
Junyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min, Liang Yang, Hongfei Lin
摘要
The widespread dissemination of toxic online posts is increasingly damaging to society. However, research on detecting toxic language in Chinese has lagged significantly due to limited datasets. Existing datasets suffer from a lack of fine-grained annotations, such as the toxic type and expressions with indirect toxicity. These fine-grained annotations are crucial factors for accurately detecting the toxicity of posts involved with lexical knowledge, which has been a challenge for researchers. To tackle this problem, we facilitate the fine-grained detection of Chinese toxic language by building a new dataset with benchmark results. First, we devised Monitor Toxic Frame, a hierarchical taxonomy to analyze the toxic type and expressions. Then, we built a fine-grained dataset ToxiCN, including both direct and indirect toxic samples. ToxiCN is based on an insulting vocabulary containing implicit profanity. We further propose a benchmark model, Toxic Knowledge Enhancement (TKE), by incorporating lexical features to detect toxic language. We demonstrate the usability of ToxiCN and the effectiveness of TKE based on a systematic quantitative and qualitative analysis.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- MultiHateClip: A Multilingual Benchmark Dataset for Hateful Video Detection on YouTube and BilibiliHan Wang, Tan Rui Yang, Usman Naseem, Roy Ka-Wei LeeACM MM 2024 · 被引用 23 次
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 被引用 5 次
- Redefining Experts: Interpretable Decomposition of Language Models for Toxicity MitigationZuhair Hasan Shaik, Abdullah Mazhar, Aseem Srivastava, Md. Shad AkhtarNeurIPS 2025 · 被引用 3 次
- MMBERT: Scaled Mixture-of-Experts Multimodal BERT for Robust Chinese Hate Speech Detection Under Cloaking PerturbationsQiyao Xue, Yuchen Dou, Zheyuan Ryan Shi, Xiang Lorraine Li 等AAAI 2026 · 被引用 2 次
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding 等EMNLP 2025 · 被引用 1 次
它引用的顶会 Paper6
- Latent Hatred: A Benchmark for Understanding Implicit Hate SpeechMai ElSherief, Caleb Ziems, David Muchlinski, Vaishnavi Anupindi 等EMNLP 2021 · 被引用 159 次
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng 等EMNLP 2022 · 被引用 82 次
- ToKen: Task Decomposition and Knowledge Infusion for Few-Shot Hate Speech DetectionBadr AlKhamissi, Faisal Ladhak, Srini Iyer, Veselin Stoyanov 等EMNLP 2022 · 被引用 17 次
- Generating Counter Narratives against Online Hate Speech: Data and StrategiesSerra Sinem Tekiroglu, Yi-Ling Chung, Marco GueriniACL 2020 · 被引用 13 次
- ToxiGen: A Large-Scale Machine-Generated Dataset for Adversarial and Implicit Hate Speech DetectionThomas Hartvigsen, Saadia Gabriel, Hamid Palangi, Maarten Sap 等ACL 2022
相关 Paper
- Towards Robust Detection of Chinese Toxic Variants via Dynamic Knowledge Graph-LLM ReasoningShaochen Yang, Kefei Zhou, Wei XuWWW 2026
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng 等EMNLP 2025
- Beyond Content: A Comprehensive Speech Toxicity Dataset and Detection Framework Incorporating Paralinguistic CuesZhongjie Ba, Liang Yi, Peng Cheng, Qingcao Li 等AAAI 2026
- TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine TranslationJinfeng Li, Tianyu Du, Shouling Ji, Rong Zhang 等USENIX Security 2020
- Culture Matters in Toxic Language Detection in PersianZahra Bokaei, Walid Magdy, Bonnie WebberACL 2025
