Breaking Free from Ivory Tower: Evaluating and Enhancing Real-world Chinese Underground Adversarial Jargon Detection
Zhifan Jiang, Mingxuan Liu, Yue Qin, Baojun Liu
摘要
Underground jargon in online ecosystems threatens platform safety and public trust by evading automated content moderation and thereby concealing criminal coordination across fraud, gambling, and illicit commerce, particularly in Chinese, given the language's unique graphophonemic properties and large user base. Despite machine learning-based content moderation, underground actors increasingly evade detection with human-crafted adversarial perturbations that differ fundamentally from budget-constrained algorithmic attacks, highlighting a research gap in understanding and debunking real-world evasion techniques. In collaboration with a leading security company, we annotate and release the first large-scale, in-the-wild dataset of adversarial Chinese underground jargon and uncover its distinct characteristics-higher perturbation intensity, greater diversity of perturbation forms, and severe structural disruption-resulting in readability degradation and tokenization collapse that exacerbate detection vulnerabilities. Our systematic measurement study across state-of-the-art large language models (LLMs) further confirms the challenges in recognizing real-world adversarial Chinese jargon: even advanced models (e. g., GPT-4o) exhibit notable limitations, achieving only 88.16% and 62.74% accuracy in jargon detection and restoration, respectively, and 76.05% accuracy in illicit content detection. To address these challenges, we propose JADE, an LLM-based detection framework grounded in a taxonomy of perturbation patterns systematically derived from annotated real-world data and designed to enable reasoning beyond explicitly observed variants. This taxonomy informs a realistic adversarial learning curriculum that combines data augmentation, external knowledge retrieval, and internal adaptation, aligning the model with real-world jargon semantics. Experiments show that JADE significantly outperforms both commercial and fine-tuned open-source LLMs, achieving 98.59%, 95.91% accuracy in jargon detection and restoration, and 97.99% accuracy in illicit content detection. Moreover, it generalizes well to other downstream applications threatened by adversarial text, with F1 scores of 93.25% and 94.33% on harmful text and fraud detection, respectively. Overall, our work takes a first step toward leveraging LLM reasoning and structured perturbation knowledge to defend against real-world adversarial jargon.
问问这篇 Paper
问问你的智能体。
Lune 读过与它相关的顶会 Paper,每个回答都会注明依据哪几篇。
相关 Paper
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng 等EMNLP 2025
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 被引用 5 次
- Into the Gray Zone: Domain Contexts Can Blur LLM Safety BoundariesKi Sen Hung, Xi Yang, Chang Liu, Haoran Li 等ACL 2026 · 被引用 1 次
- "Do Anything Now": Characterizing and Evaluating In-The-Wild Jailbreak Prompts on Large Language ModelsXinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen 等CCS 2024 · 被引用 132 次
- Scalable Detection of Promotional Website Defacements in Black Hat SEO CampaignsRonghai Yang, Xianbo Wang, Cheng Chi, Dawei Wang 等USENIX Security 2021 · 被引用 27 次
