Adversarial Examples Are Not Real Features
Ang Li, Yifei Wang, Yiwen Guo, Yisen Wang
Abstract
The existence of adversarial examples has been a mystery for years and attracted much interest. A well-known theory by Ilyas et al. [18] explains adversarial vulnerability from a data perspective by showing that one can extract non-robust features from adversarial examples and these features alone are useful for classification. However, the explanation remains quite counter-intuitive since non-robust features are mostly noise features to humans. In this paper, we re-examine the theory from a larger context by incorporating multiple learning paradigms. Notably, we find that contrary to their good usefulness under supervised learning, nonrobust features attain poor usefulness when transferred to other self-supervised learning paradigms, such as contrastive learning, masked image modeling, and diffusion models. It reveals that non-robust features are not really as useful as robust or natural features that enjoy good transferability between these paradigms. Meanwhile, for robustness, we also show that naturally trained self-supervised encoders from robust features are largely non-robust under AutoAttack. Our crossparadigm examination suggests that the non-robust features are not really useful but more like paradigm-wise shortcuts, and robust features alone might be insufficient to attain reliable model robustness across paradigms. Code is available at https://github.com/PKU-ML/AdvNotRealFeatures .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 9194e6c2-7324-4eb8-bacc-ed01d7a3ff13Cited by top-tier papers5
- Balance, Imbalance, and Rebalance: Understanding Robust Overfitting from a Minimax Game PerspectiveYifei Wang, Liangchen Li, Jiansheng Yang, Zhouchen Lin et al.NeurIPS 2023 · 26 citations
- PID: Prompt-Independent Data Protection Against Latent Diffusion ModelsAng Li, Yichuan Mo, Mingjie Li, Yisen WangICML 2024 · 5 citations
- Adversarial Vulnerability from Interference Between Features in SuperpositionEdward Stevinson, Lucas Prieto, Melih Barsbey, Tolga BirdalICML 2026 · 4 citations
- Feature compression is the root cause of adversarial fragility in neural networksJingchao Gao, Ziqing Lu, Raghu Mudumbai, Xiaodong Wu et al.ICLR 2026 · 3 citations
- Identifying and Understanding Cross-Class Features in Adversarial TrainingZeming Wei, Steven Y. Guo, Yisen WangICML 2025
Builds on27
- Denoising Diffusion Probabilistic ModelsJonathan Ho, Ajay Jain, Pieter AbbeelNeurIPS 2020 · 35,902 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- BEiT: BERT Pre-Training of Image TransformersHangbo Bao, Li Dong, Songhao Piao, Furu WeiICLR 2022 · 3,632 citations
- Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacksFrancesco Croce, Matthias HeinICML 2020 · 2,337 citations
Related papers
- Self-Supervised Adversarial Training via Diverse Augmented Queries and Self-Supervised Double PerturbationRuize Zhang, Sheng Tang, Juan CaoNeurIPS 2024 · 5 citations
- Boosting Generative Adversarial Transferability with Self-Supervised Vision Transformer FeaturesShangbo Wu, Yu-an Tan, Ruinan Ma, Wencong Ma et al.ICCV 2025
- When does Contrastive Learning Preserve Adversarial Robustness from Pretraining to Finetuning?Lijie Fan, Sijia Liu, Pin-Yu Chen, Gaoyuan Zhang et al.NeurIPS 2021 · 147 citations
- ArCL: Enhancing Contrastive Learning with Augmentation-Robust RepresentationsXuyang Zhao, Tianqi Du, Yisen Wang, Jun Yao et al.ICLR 2023 · 2 citations
- Beyond Pretrained Features: Noisy Image Modeling Provides Adversarial DefenseZunzhi You, Daochang Liu, Bohyung Han, Chang XuNeurIPS 2023 · 9 citations
