Adversarial Examples Are Not Real Features
Ang Li, Yifei Wang, Yiwen Guo, Yisen Wang
摘要
The existence of adversarial examples has been a mystery for years and attracted much interest. A well-known theory by Ilyas et al. [18] explains adversarial vulnerability from a data perspective by showing that one can extract non-robust features from adversarial examples and these features alone are useful for classification. However, the explanation remains quite counter-intuitive since non-robust features are mostly noise features to humans. In this paper, we re-examine the theory from a larger context by incorporating multiple learning paradigms. Notably, we find that contrary to their good usefulness under supervised learning, nonrobust features attain poor usefulness when transferred to other self-supervised learning paradigms, such as contrastive learning, masked image modeling, and diffusion models. It reveals that non-robust features are not really as useful as robust or natural features that enjoy good transferability between these paradigms. Meanwhile, for robustness, we also show that naturally trained self-supervised encoders from robust features are largely non-robust under AutoAttack. Our crossparadigm examination suggests that the non-robust features are not really useful but more like paradigm-wise shortcuts, and robust features alone might be insufficient to attain reliable model robustness across paradigms. Code is available at https://github.com/PKU-ML/AdvNotRealFeatures .
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper5
- Balance, Imbalance, and Rebalance: Understanding Robust Overfitting from a Minimax Game PerspectiveYifei Wang, Liangchen Li, Jiansheng Yang, Zhouchen Lin 等NeurIPS 2023 · 被引用 26 次
- PID: Prompt-Independent Data Protection Against Latent Diffusion ModelsAng Li, Yichuan Mo, Mingjie Li, Yisen WangICML 2024 · 被引用 5 次
- Adversarial Vulnerability from Interference Between Features in SuperpositionEdward Stevinson, Lucas Prieto, Melih Barsbey, Tolga BirdalICML 2026 · 被引用 4 次
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu 等ICLR 2026 · 被引用 3 次
- Identifying and Understanding Cross-Class Features in Adversarial TrainingZeming Wei, Steven Y. Guo, Yisen WangICML 2025
它引用的顶会 Paper27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 被引用 35,902 次
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 被引用 3,632 次
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 被引用 2,337 次
相关 Paper
- Self-Supervised Adversarial Training via Diverse Augmented Queries and Self-Supervised Double PerturbationRuize Zhang, Sheng Tang, Juan CaoNeurIPS 2024 · 被引用 5 次
- Boosting Generative Adversarial Transferability with Self-Supervised Vision Transformer FeaturesShangbo Wu, Yu-an Tan, Ruinan Ma, Wencong Ma 等ICCV 2025
- When does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?Lijie Fan, Sijia Liu, Pin-Yu Chen, Gaoyuan Zhang 等NeurIPS 2021 · 被引用 147 次
- ArCL: Enhancing Contrastive Learning with Augmentation-Robust RepresentationsXuyang Zhao, Tianqi Du, Yisen Wang, Jun Yao 等ICLR 2023 · 被引用 2 次
- Beyond Pretrained Features: Noisy Image Modeling Provides Adversarial DefenseZunzhi You, Daochang Liu, Bohyung Han, Chang XuNeurIPS 2023 · 被引用 9 次
