Adversarially Pretrained Transformers May Be Universally Robust In-Context Learners
Soichiro Kumano, Hiroshi Kera, Toshihiko Yamasaki
摘要
Adversarial training is one of the most effective defenses against adversarial attacks, but it incurs a high computational cost. In this study, we present the first theoretical analysis suggesting that adversarially pretrained transformers can serve as universally robust foundation models-models that can adapt robustly to diverse downstream tasks with only lightweight tuning. Specifically, we demonstrate that single-layer linear transformers, after adversarial pretraining across a variety of classification tasks, can generalize robustly to unseen classification tasks through in-context learning from clean demonstrations (i.e., without requiring additional adversarial training or examples). This universal robustness stems from the model's ability to adaptively focus on robust features within given tasks. We also identify two open challenges for attaining robustness: the accuracy-robustness trade-off and sample-hungry training. This study initiates the discussion on the utility of universally robust foundation models. While their training is expensive, the investment would prove worthwhile as downstream tasks can obtain adversarial robustness for free. The code is available at https://github.com/s-kumano/univer sally-robust-in-context-learner .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Benign Overfitting in Adversarial Training for Vision TransformersJiaming Zhang, Meng Ding, Shaopeng Fu, Jingfeng Zhang 等ICML 2026 · 被引用 1 次
- Understanding and Improving Continuous LLM Adversarial Training via In-context Learning TheoryShaopeng Fu, Di WangICLR 2026 · 被引用 1 次
它引用的顶会 Paper50
- Language Models are Few-Shot LearnersTom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah 等NeurIPS 2020 · 被引用 64,255 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Extracting Training Data from Large Language ModelsNicholas Carlini, Florian Tramèr, Eric Wallace, Matthew Jagielski 等USENIX Security 2021 · 被引用 2,866 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
- Fast is better than free: Revisiting adversarial trainingEric Wong, Leslie Rice, J. Zico KolterICLR 2020 · 被引用 1,352 次
相关 Paper
- Are Transformers more robust than CNNs?Yutong Bai, Jieru Mei, Alan L. Yuille, Cihang XieNeurIPS 2021 · 被引用 365 次
- Downstream-agnostic Adversarial ExamplesZiqi Zhou, Shengshan Hu, Ruizhi Zhao, Qian Wang 等ICCV 2023 · 被引用 45 次
- Adversarial Robustness: From Self-Supervised Pre-Training to Fine-TuningTianlong Chen, Sijia Liu, Shiyu Chang, Yu Cheng 等CVPR 2020
- Short-length Adversarial Training Helps LLMs Defend Long-length Jailbreak Attacks: Theoretical and Empirical EvidenceShaopeng Fu, Liang Ding, Jingfeng Zhang, Di WangNeurIPS 2025 · 被引用 15 次
- Adversarial Self-Attention for Language UnderstandingHongqiu Wu, Ruixue Ding, Hai Zhao, Pengjun Xie 等AAAI 2023 · 被引用 19 次
