How Should Pre-Trained Language Models Be Fine-Tuned Towards Adversarial Robustness?
Xinshuai Dong, Anh Tuan Luu, Min Lin, Shuicheng Yan, Hanwang Zhang
摘要
The fine-tuning of pre-trained language models has a great success in many NLP fields. Yet, it is strikingly vulnerable to adversarial examples, e.g., word substitution attacks using only synonyms can easily fool a BERT-based sentiment analysis model. In this paper, we demonstrate that adversarial training, the prevalent defense technique, does not directly fit a conventional fine-tuning scenario, because it suffers severely from catastrophic forgetting: failing to retain the generic and robust linguistic features that have already been captured by the pre-trained model. In this light, we propose Robust Informative Fine-Tuning (RIFT), a novel adversarial fine-tuning method from an information-theoretical perspective. In particular, RIFT encourages an objective model to retain the features learned from the pre-trained model throughout the entire fine-tuning process, whereas a conventional one only uses the pre-trained weights for initialization. Experimental results show that RIFT consistently outperforms the state-of-the-arts on two popular NLP tasks: sentiment analysis and natural language inference, under different attacks across various pre-trained language models. 1 .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- Model Tailor: Mitigating Catastrophic Forgetting in Multi-modal Large Language ModelsDidi Zhu, Zhongyi Sun, Zexi Li, Tao Shen 等ICML 2024 · 被引用 50 次
- Certified Robustness Against Natural Language Attacks by Causal InterventionHaiteng Zhao, Chang Ma, Xinshuai Dong, Anh Tuan Luu 等ICML 2022 · 被引用 43 次
- Prompt as Triggers for Backdoor Attack: Examining the Vulnerability in Language ModelsShuai Zhao, Jinming Wen, Anh Tuan Luu, Junbo Zhao 等EMNLP 2023 · 被引用 39 次
- Pre-trained Adversarial PerturbationsYuanhao Ban, Yinpeng DongNeurIPS 2022 · 被引用 37 次
- InfoCTM: A Mutual Information Maximization Perspective of Cross-Lingual Topic ModelingXiaobao Wu, Xinshuai Dong, Thong Nguyen, Chaoqun Liu 等AAAI 2023 · 被引用 35 次
它引用的顶会 Paper14
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Towards Evaluating the Robustness of Neural NetworksNicholas Carlini, David A. WagnerS&P 2017 · 被引用 9,786 次
- Understanding Contrastive Representation Learning through Alignment and Uniformity on the HypersphereTongzhou Wang, Phillip IsolaICML 2020 · 被引用 2,360 次
- What Makes for Good Views for Contrastive Learning?Yonglong Tian, Chen Sun, Ben Poole, Dilip Krishnan 等NeurIPS 2020 · 被引用 1,631 次
- Contrastive Representation DistillationYonglong Tian, Dilip Krishnan, Phillip IsolaICLR 2020 · 被引用 1,305 次
相关 Paper
- InfoBERT: Improving Robustness of Language Models from An Information Theoretic PerspectiveBoxin Wang, Shuohang Wang, Yu Cheng, Zhe Gan 等ICLR 2021 · 被引用 132 次
- ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsLan Jiang, Hao Zhou, Yankai Lin, Peng Li 等EMNLP 2022 · 被引用 5 次
- Model-tuning Via Prompts Makes NLP Models Adversarially RobustMrigank Raman, Pratyush Maini, J. Zico Kolter, Zachary C. Lipton 等EMNLP 2023 · 被引用 7 次
- Improving Generalization of Adversarial Training via Robust Critical Fine-TuningKaijie Zhu, Xixu Hu, Jindong Wang, Xing Xie 等ICCV 2023 · 被引用 38 次
- Pre-Trained Model Guided Fine-Tuning for Zero-Shot Adversarial RobustnessSibo Wang, Jie Zhang, Zheng Yuan, Shiguang ShanCVPR 2024
