SHIELD: Defending Textual Neural Networks against Multiple Black-Box Adversarial Attacks with Stochastic Multi-Expert Patcher
Thai Le, Noseong Park, Dongwon Lee
摘要
Even though several methods have proposed to defend textual neural network (NN) models against black-box adversarial attacks, they often defend against a specific text perturbation strategy and/or require re-training the models from scratch. This leads to a lack of generalization in practice and redundant computation. In particular, the state-of-the-art transformer models (e.g., BERT, RoBERTa) require great time and computation resources. By borrowing an idea from software engineering, in order to address these limitations, we propose a novel algorithm, SHIELD, which modifies and re-trains only the last layer of a textual NN, and thus it “patches” and “transforms” the NN into a stochastic weighted ensemble of multi-expert prediction heads. Considering that most of current black-box attacks rely on iterative search mechanisms to optimize their adversarial perturbations, SHIELD confuses the attackers by automatically utilizing different weighted ensembles of predictors depending on the input. In other words, SHIELD breaks a fundamental assumption of the attack, which is a victim NN model remains constant during an attack. By conducting comprehensive experiments, we demonstrate that all of CNN, RNN, BERT, and RoBERTa-based textual NNs, once patched by SHIELD, exhibit a relative enhancement of 15%–70% in accuracy on average against 14 different black-box attacks, outperforming 6 defensive baselines across 3 public datasets. All codes are to be released.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了最后一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper6
- HQA-Attack: Toward High Quality Black-Box Hard-Label Adversarial Attack on TextHan Liu, Zhi Xu, Xiaotong Zhang, Feng Zhang 等NeurIPS 2023 · 被引用 32 次
- RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial AttacksZhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su 等ACL 2023 · 被引用 16 次
- Don't Retrain, Just Rewrite: Countering Adversarial Perturbations by Rewriting TextAshim Gupta, Carter Wood Blum, Temma Choji, Yingjie Fei 等ACL 2023 · 被引用 7 次
- ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsLan Jiang, Hao Zhou, Yankai Lin, Peng Li 等EMNLP 2022 · 被引用 5 次
- TextShield: Beyond Successfully Detecting Adversarial Sentences in text classificationLingfeng Shen, Ze Zhang, Haiyun Jiang, Ying ChenICLR 2023
它引用的顶会 Paper9
- Generative Pretraining From PixelsMark Chen, Alec Radford, Rewon Child, Jeffrey Wu 等ICML 2020 · 被引用 1,773 次
- Word-level Textual Adversarial Attacking as Combinatorial OptimizationYuan Zang, Fanchao Qi, Chenghao Yang, Zhiyuan Liu 等ACL 2020 · 被引用 188 次
- Mixup Inference: Better Exploiting Mixup to Defend Adversarial AttacksTianyu Pang, Kun Xu, Jun ZhuICLR 2020 · 被引用 114 次
- Adversarial Training with Fast Gradient Projection Method against Synonym Substitution Based Text AttacksXiaosen Wang, Yichen Yang, Yihe Deng, Kun HeAAAI 2021 · 被引用 98 次
- Towards Robustness Against Natural Language Word SubstitutionsXinshuai Dong, Anh Tuan Luu, Rongrong Ji, Hong LiuICLR 2021 · 被引用 63 次
相关 Paper
- Improving the Robustness of Transformer-based Large Language Models with Dynamic AttentionLujia Shen, Yuwen Pu, Shouling Ji, Changjiang Li 等NDSS 2024
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li 等EMNLP 2021 · 被引用 46 次
- Fooling Explanations in Text ClassifiersAdam Ivankay, Ivan Girardi, Chiara Marchiori, Pascal FrossardICLR 2022 · 被引用 23 次
- ProTransformer: Robustify Transformers via Plug-and-Play ParadigmZhichao Hou, Weizhi Gao, Yuchen Shen, Feiyi Wang 等NeurIPS 2024 · 被引用 9 次
- Text Adversarial Purification as Defense against Adversarial AttacksLinyang Li, Demin Song, Xipeng QiuACL 2023 · 被引用 11 次
