Word Level Robustness Enhancement: Fight Perturbation with Perturbation
Pei Huang, Yuting Yang, Fuqi Jia, Minghao Liu, Feifei Ma, Jian Zhang
摘要
State-of-the-art deep NLP models have achieved impressive improvements on many tasks. However, they are found to be vulnerable to some perturbations. Before they are widely adopted, the fundamental issues of robustness need to be addressed. In this paper, we design a robustness enhancement method to defend against word substitution perturbation, whose basic idea is to fight perturbation with perturbation. We find that: although many well-trained deep models are not robust in the setting of the presence of adversarial samples, they satisfy weak robustness. That means they can handle most non-crafted perturbations well. Taking advantage of the weak robustness property of deep models, we utilize non-crafted perturbations to resist the adversarial perturbations crafted by attackers. Our method contains two main stages. The first stage is using randomized perturbation to conform the input to the data distribution. The second stage is using randomized perturbation to eliminate the instability of prediction results and enhance the robustness guarantee. Experimental results show that our method can significantly improve the ability of deep models to resist the state-of-the-art adversarial attacks while maintaining the prediction performance on the original clean data.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper4
- Text-CRS: A Generalized Certified Robustness Framework against Textual Adversarial AttacksXinyu Zhang, Hanbin Hong, Yuan Hong, Peng Huang 等S&P 2024 · 被引用 41 次
- Towards the Robustness of Differentially Private Federated LearningTao Qi, Huili Wang, Yongfeng HuangAAAI 2024 · 被引用 30 次
- CaDRec: Contextualized and Debiased Recommender ModelXinfeng Wang, Fumiyo Fukumoto, Jin Cui, Yoshimi Suzuki 等SIGIR 2024 · 被引用 4 次
- Breaking the Adversarial Robustness-Performance Trade-off in Text Classification via Manifold PurificationChenhao Dang, Jing MaAAAI 2026
它引用的顶会 Paper10
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Automatic Perturbation Analysis for Scalable Certified Robustness and BeyondKaidi Xu, Zhouxing Shi, Huan Zhang, Yihan Wang 等NeurIPS 2020 · 被引用 415 次
- Seq2Sick: Evaluating the Robustness of Sequence-to-Sequence Models with Adversarial ExamplesMinhao Cheng, Jinfeng Yi, Pin-Yu Chen, Huan Zhang 等AAAI 2020 · 被引用 268 次
- Understanding and Mitigating the Tradeoff between Robustness and AccuracyAditi Raghunathan, Sang Michael Xie, Fanny Yang, John C. Duchi 等ICML 2020 · 被引用 252 次
相关 Paper
- Searching for an Effective Defender: Benchmarking Defense against Adversarial Word SubstitutionZongyi Li, Jianhan Xu, Jiehang Zeng, Linyang Li 等EMNLP 2021 · 被引用 46 次
- Defense against Synonym Substitution-based Adversarial Attacks via Dirichlet Neighborhood EnsembleYi Zhou, Xiaoqing Zheng, Cho-Jui Hsieh, Kai-Wei Chang 等ACL 2021
- Robustness to Programmable String Transformations via Augmented Abstract TrainingYuhao Zhang, Aws Albarghouthi, Loris D'AntoniICML 2020 · 被引用 17 次
- Towards Robustness Against Natural Language Word SubstitutionsXinshuai Dong, Anh Tuan Luu, Rongrong Ji, Hong LiuICLR 2021 · 被引用 63 次
- ε-weakened robustness of deep neural networksPei Huang, Yuting Yang, Minghao Liu, Fuqi Jia 等ISSTA 2022 · 被引用 10 次
