Balanced Adversarial Training: Balancing Tradeoffs between Fickleness and Obstinacy in NLP Models
Hannah Chen, Yangfeng Ji, David E. Evans
摘要
Traditional (fickle) adversarial examples involve finding a small perturbation that does not change an input's true label but confuses the classifier into outputting a different prediction. Conversely, obstinate adversarial examples occur when an adversary finds a small perturbation that preserves the classifier's prediction but changes the true label of an input. Adversarial training and certified robust training have shown some effectiveness in improving the robustness of machine learnt models to fickle adversarial examples. We show that standard adversarial training methods focused on reducing vulnerability to fickle adversarial examples may make a model more vulnerable to obstinate adversarial examples, with experiments for both natural language inference and paraphrase identification tasks. To counter this phenomenon, we introduce Balanced Adversarial Training, which incorporates contrastive learning to increase robustness against both fickle and obstinate adversarial examples.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper6
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 被引用 1,333 次
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue 等EMNLP 2020 · 被引用 529 次
- Adversarial Self-Supervised Contrastive LearningMinseon Kim, Jihoon Tack, Sung Ju HwangNeurIPS 2020 · 被引用 294 次
- Contrastive Learning with Adversarial ExamplesChih-Hui Ho, Nuno VasconcelosNeurIPS 2020 · 被引用 174 次
相关 Paper
- Fundamental Tradeoffs between Invariance and Sensitivity to Adversarial PerturbationsFlorian Tramèr, Jens Behrmann, Nicholas Carlini, Nicolas Papernot 等ICML 2020 · 被引用 103 次
- Contrast and Classify: Training Robust VQA ModelsYash Kant, Abhinav Moudgil, Dhruv Batra, Devi Parikh 等ICCV 2021 · 被引用 30 次
- Adversarial Invariant LearningNanyang Ye, Jingxuan Tang, Huayu Deng, Xiao-Yun Zhou 等CVPR 2021
- On the Vulnerability of Adversarially Trained Models Against Two-faced AttacksShengjie Zhou, Lue Tao, Yuzhou Cao, Tao Xiang 等ICLR 2024
- Robust Pre-Training by Adversarial Contrastive LearningZiyu Jiang, Tianlong Chen, Ting Chen, Zhangyang WangNeurIPS 2020 · 被引用 284 次
