How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, Hanwang Zhang
Abstract
The fine-tuning of pre-trained language models has a great success in many NLP fields. Yet, it is strikingly vulnerable to adversarial examples, e.g., word substitution attacks using only synonyms can easily fool a BERT-based sentiment analysis model. In this paper, we demonstrate that adversarial training, the prevalent defense technique, does not directly fit a conventional fine-tuning scenario, because it suffers severely from catastrophic forgetting: failing to retain the generic and robust linguistic features that have already been captured by the pre-trained model. In this light, we propose Robust Informative Fine-Tuning (RIFT), a novel adversarial fine-tuning method from an information-theoretical perspective. In particular, RIFT encourages an objective model to retain the features learned from the pre-trained model throughout the entire fine-tuning process, whereas a conventional one only uses the pre-trained weights for initialization. Experimental results show that RIFT consistently outperforms the state-of-the-arts on two popular NLP tasks: sentiment analysis and natural language inference, under different attacks across various pre-trained language models. 1 .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext b6c87d18-b8bd-43b0-901f-4f63785fae35Cited by top-tier papers13
- Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language ModelsDidi Zhu, Zhongyi Sun, Zexi Li, Tao Shen et al.ICML 2024 · 50 citations
- Certified Robustness Against Natural Language Attacks by Causal InterventionHaiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu et al.ICML 2022 · 43 citations
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsShuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao et al.EMNLP 2023 · 39 citations
- Pre-trained Adversarial PerturbationsYuanhao Ban, Yinpeng DongNeurIPS 2022 · 37 citations
- InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic ModelingXiaobao Wu, Xinshuai Dong, Thong Nguyen, Chaoqun Liu et al.AAAI 2023 · 35 citations
Builds on14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 9,786 citations
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 2,360 citations
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan et al.NeurIPS 2020 · 1,631 citations
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 1,305 citations
Related papers
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan et al.ICLR 2021 · 132 citations
- ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsLan Jiang, Hao Zhou, Yankai Lin, Peng Li et al.EMNLP 2022 · 5 citations
- Model-tuning Via Prompts Makes NLP Models Adversarially RobustMrigank Raman, Pratyush Maini, J. Zico Kolter, Zachary C. Lipton et al.EMNLP 2023 · 7 citations
- Improving Generalization of Adversarial Training via Robust Critical Fine-TuningKaijie Zhu, Xixu Hu, Jindong Wang, Xing Xie et al.ICCV 2023 · 38 citations
- Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial RobustnessSibo Wang, Jie Zhang, Zheng Yuan, Shiguang ShanCVPR 2024
