New Terms, New Toxicity: Consensus-based Chinese Neologism Toxicity Detection via Search-Augmented LLMs
Shiyao Cui, Qinglin Zhang, Di Wang, Yida Lu, Zhexin Zhang, Jinhua Gao, Jinglin Yang, Min He, Han Qiu, Minlie Huang
Abstract
Neologisms, emerging terms in meaning or form, can serve as new vehicles for toxic expression, like "田 园 女" ("country girl") as a stigmatizing label targeting feminism. Such toxic neologisms appear benign but have evolved into toxic usage in public consensus, posing challenges to moderation systems and remaining underexplored. In this paper, we investigate how to detect implicit toxicity expressed via neologisms. We first propose a taxonomy that captures the origins and consensusverification criteria of toxic neologisms, followed by the construction of a lexicon spanning widely observed risk categories. To capture toxicity grounded in public consensus, we introduce SeTox, a search-augmented framework that enables static large language models (LLMs) to incorporate real-time web context for neologism toxicity detection. Experiments show that SeTox, even with 3B-scale models, outperforms recent large-scale models, revealing its scalability to incorporate realworld knowledge for toxic neologism detection. Disclaimer: this paper has contents that may be disturbing to some readers.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 92629ff5-e60f-4123-9ac2-8f4139034becBuilds on8
- COLD: A Benchmark for Chinese Offensive Language DetectionJiawen Deng, Jingyan Zhou, Hao Sun, Chujie Zheng et al.EMNLP 2022 · 82 citations
- Facilitating Fine-grained Detection of Chinese Toxic Language: Hierarchical Taxonomy, Resources, and BenchmarksJunyu Lu, Bo Xu, Xiaokun Zhang, Changrong Min et al.ACL 2023 · 25 citations
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 5 citations
- Sheep's Skin, Wolf's Deeds: Are LLMs Ready for Metaphorical Implicit Hate Speech?Jingjie Zeng, Liang Yang, Zekun Wang, Yuanyuan Sun et al.ACL 2025 · 4 citations
- Speculating LLMs' Chinese Training Data Pollution from Their TokensQingjie Zhang, Di Wang, Haoting Qian, Liu Yan et al.EMNLP 2025 · 1 citation
Related papers
- NEO-BENCH: Evaluating Robustness of Large Language Models with NeologismsJonathan Zheng, Alan Ritter, Wei XuACL 2024
- Pragmatic Inference Chain (PIC) Improving LLMs' Reasoning of Authentic Implicit Toxic LanguageXi Chen, Shuo WangEMNLP 2025 · 7 citations
- Toxicity Red-Teaming: Benchmarking LLM Safety in Singapore's Low-Resource LanguagesYujia Hu, Ming Shan Hee, Preslav Nakov, Roy Ka-Wei LeeEMNLP 2025
- Towards Robust Detection of Chinese Toxic Variants via Dynamic Knowledge Graph-LLM ReasoningShaochen Yang, Kefei Zhou, Wei XuWWW 2026
- ModelCitizens: Representing Community Voices in Online SafetyAshima Suvarna, Christina Chance, Karolina Naranjo, Hamid Palangi et al.EMNLP 2025
