Flooding-X: Improving BERT's Resistance to Adversarial Attacks via Loss-Restricted Fine-Tuning
Qin Liu, Rui Zheng, Bao Rong, Jingyi Liu, Zhihua Liu, Zhanzhan Cheng, Liang Qiao, Tao Gui, Qi Zhang, Xuanjing Huang
Abstract
Adversarial robustness has attracted much attention recently, and the mainstream solution is adversarial training. However, the tradition of generating adversarial perturbations for each input embedding (in the settings of NLP) scales up the training computational complexity by the number of gradient steps it takes to obtain the adversarial samples. To address this problem, we leverage Flooding method which primarily aims at better generalization and we find promising in defending adversarial attacks. We further propose an effective criterion to bring hyper-parameter-dependent flooding into effect with a narrowed-down search space by measuring how the gradient steps taken within one epoch affect the loss of each batch. Our approach requires zero adversarial sample for training, and its time consumption is equivalent to fine-tuning, which can be 2-15 times faster than standard adversarial training. We experimentally show that our method improves BERT's resistance to textual adversarial attacks by a large margin, and achieves state-of-the-art robust accuracy on various text classification and GLUE tasks. * * Equal Contribution.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 746302ef-de70-4734-b253-cb7e292c7973Cited by top-tier papers10
- RMLM: A Flexible Defense Framework for Proactively Mitigating Word-level Adversarial AttacksZhaoyang Wang, Zhiyue Liu, Xiaopeng Zheng, Qinliang Su et al.ACL 2023 · 16 citations
- Multi-CLS BERT: An Efficient Alternative to Traditional EnsemblingHaw-Shiuan Chang, Ruei-Yao Sun, Kathryn Ricci, Andrew McCallumACL 2023 · 8 citations
- Robust Few-Shot Named Entity Recognition with Boundary Discrimination and Correlation PurificationXiaojun Xue, Chunxia Zhang, Tianxiang Xu, Zhendong NiuAAAI 2024 · 7 citations
- Efficient Adversarial Training with Robust Early-Bird TicketsZhiheng Xi, Rui Zheng, Tao Gui, Qi Zhang et al.EMNLP 2022 · 7 citations
- Rethinking Consistent Multi-Label Classification Under Inexact SupervisionWei Wang, Tianhao Ma, Ming-Kun Xie, Gang Niu et al.ICLR 2026 · 3 citations
Builds on13
- Is BERT Really Robust? A Strong Baseline for Natural Language Attack on Text Classification and EntailmentDi Jin, Zhijing Jin, Joey Tianyi Zhou, Peter SzolovitsAAAI 2020 · 1,333 citations
- TextBugger: Generating Adversarial Text Against Real-world ApplicationsJinfeng Li, Shouling Ji, Tianyu Du, Bo Li et al.NDSS 2019 · 876 citations
- Adversarial NLI: A New Benchmark for Natural Language UnderstandingYixin Nie, Adina Williams, Emily Dinan, Mohit Bansal et al.ACL 2020 · 602 citations
- BERT-ATTACK: Adversarial Attack Against BERT Using BERTLinyang Li, Ruotian Ma, Qipeng Guo, Xiangyang Xue et al.EMNLP 2020 · 529 citations
- FreeLB: Enhanced Adversarial Training for Natural Language UnderstandingChen Zhu, Yu Cheng, Zhe Gan, Siqi Sun et al.ICLR 2020 · 502 citations
Related papers
- DSRM: Boost Textual Adversarial Training with Distribution Shift Risk MinimizationSongyang Gao, Shihan Dou, Yan Liu, Xiao Wang et al.ACL 2023 · 3 citations
- Adversarial Training with Fast Gradient Projection Method against Synonym Substitution Based Text AttacksXiaosen Wang, Yichen Yang, Yihe Deng, Kun HeAAAI 2021 · 98 citations
- Improving the Robustness of Transformer-based Large Language Models with Dynamic AttentionLujia Shen, Yuwen Pu, Shouling Ji, Changjiang Li et al.NDSS 2024
- Generative Adversarial Training with Perturbed Token Detection for Model RobustnessJiahao Zhao, Wenji MaoEMNLP 2023 · 3 citations
- Token-Aware Virtual Adversarial Training in Natural Language UnderstandingLinyang Li, Xipeng QiuAAAI 2021 · 54 citations
