USENIX Security2020Top-tier venue
TextShield: Robust Text Classification Based on Multimodal Embedding and Neural Machine Translation
Jinfeng Li, Tianyu Du, Shouling Ji, Rong Zhang, Quan Lu, Min Yang, Ting Wang
Abstract
Text-based toxic content detection is an important tool for reducing harmful interactions in online social media environments. Yet, its underlying mechanism, deep learning-based text classification (DLTC), is inherently vulnerable to maliciously crafted adversarial texts. To mitigate such vulnerabilities, intensive research has been conducted on strengthening English-based DLTC models. However, the existing defenses are not effective for Chinese-based DLTC models, due to the unique sparseness, diversity, and variation of the Chinese language. In this paper, we bridge this striking gap by presenting TEXTSHIELD, a new adversarial defense framework specifically designed for Chinese-based DLTC models. TEXTSHIELD differs from previous work in several key aspects: (i) generic -it applies to any Chinese-based DLTC models without requiring re-training; (ii) robust -it significantly reduces the attack success rate even under the setting of adaptive attacks; and (iii) accurate -it has little impact on the performance of DLTC models over legitimate inputs. Extensive evaluations show that it outperforms both existing methods and the industry-leading platforms. Future work will explore its applicability in broader practical tasks.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext eb4c5c8f-a866-44da-be00-48e560cf023bCited by top-tier papers6
- MalGraph: Hierarchical Graph Neural Networks for Robust Windows Malware DetectionXiang Ling, Lingfei Wu, Wei Deng, Zhenqing Qu et al.INFOCOM 2022 · 47 citations
- RoCBert: Robust Chinese Bert with Multimodal Contrastive PretrainingHui Su, Weiwei Shi, Xiaoyu Shen, Xiao Zhou et al.ACL 2022 · 38 citations
- "Get in Researchers; We're Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security ConferencesDaniel Olszewski, Allison Lu, Carson Stillman, Kevin Warren et al.CCS 2023 · 19 citations
- "Is your explanation stable?": A Robustness Evaluation Framework for Feature AttributionYuyou Gan, Yuhao Mao, Xuhong Zhang, Shouling Ji et al.CCS 2022 · 13 citations
- Improving the Robustness of Transformer-based Large Language Models with Dynamic AttentionLujia Shen, Yuwen Pu, Shouling Ji, Changjiang Li et al.NDSS 2024
Builds on5
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Model-Reuse Attacks on Deep Learning SystemsYujie Ji, Xinyang Zhang, Shouling Ji, Xiapu Luo et al.CCS 2018 · 197 citations
- DEEPSEC: A Uniform Platform for Security Analysis of Deep Learning ModelXiang Ling, Shouling Ji, Jiaxu Zou, Jiannan Wang et al.S&P 2019 · 147 citations
- Stealthy Porn: Understanding Real-World Adversarial Images for Illicit Online PromotionKan Yuan, Di Tang, Xiaojing Liao, XiaoFeng Wang et al.S&P 2019 · 49 citations
- Interpretable Deep Learning under FireXinyang Zhang, Ningfei Wang, Hua Shen, Shouling Ji et al.USENIX Security 2020
Related papers
- TextShield: Beyond Successfully Detecting Adversarial Sentences in text classificationLingfeng Shen, Ze Zhang, Haiyun Jiang, Ying ChenICLR 2023
- Enhancing Chinese Offensive Language Detection with Homophonic PerturbationJunqi Wu, Shujie Ji, Kang Zhong, Huiling Peng et al.EMNLP 2025
- ToxiCloakCN: Evaluating Robustness of Offensive Language Detection in Chinese with Cloaking PerturbationsYunze Xiao, Yujia Hu, Kenny T. W. Choo, Roy Ka-Wei LeeEMNLP 2024 · 5 citations
- Iron Sharpens Iron: Defending Against Attacks in Machine-Generated Text Detection with Adversarial TrainingYuanfan Li, Zhaohan Zhang, Chengzhengxu Li, Chao Shen et al.ACL 2025 · 10 citations
- Chinese Toxic Language Mitigation via Sentiment Polarity Consistent RewritesXintong Wang, Yixiao Liu, Jingheng Pan, Liang Ding et al.EMNLP 2025 · 1 citation
