Adversarial Vulnerability from Interference Between Features in Superposition
Edward Stevinson, Lucas Prieto, Melih Barsbey, Tolga Birdal
Abstract
Why do adversarial examples exist, and why do they transfer between models? Existing explanations appeal to high-dimensional geometry, non-robust patterns in the input, and decision boundary structure, but none provides a representation-level mechanism that explains why specific perturbations succeed and why attacks transfer between models. In this paper, we show that adversarial vulnerability can stem from efficient information encoding in neural networks. Specifically, vulnerability can arise from superposition - the phenomenon where networks represent more concepts than they have dimensions, forcing non-orthogonal representation and thus interference. This interference causes perturbations targeting one representation to affect others, creating vulnerabilities determined by interference patterns. In synthetic settings with precisely controlled superposition, we establish that superposition suffices to create adversarial vulnerability. The resulting attacks are predictable: PGD-discovered perturbations align with theoretically optimal perturbations derived from the interference geometry. Models trained on similar data develop similar interference patterns, explaining attack transferability. We then show that successful attacks on image classifiers exhibit the structure predicted by our proposed mechanism. These findings reveal that adversarial vulnerability can be a byproduct of networks' representational compression, complementing existing explanations based on data properties or architectural factors.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext e162a788-d251-4372-a093-04634d7f2ee8Builds on15
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- The Linear Representation Hypothesis and the Geometry of Large Language ModelsKiho Park, Yo Joong Choe, Victor VeitchICML 2024 · 461 citations
- Rank Diminishing in Deep Neural NetworksRuili Feng, Kecheng Zheng, Yukun Huang, Deli Zhao et al.NeurIPS 2022 · 64 citations
- Adversarial Robustness Limits via Scaling-Law and Human-Alignment StudiesBrian R. Bartoldson, James Diffenderfer, Konstantinos Parasyris, Bhavya KailkhuraICML 2024 · 45 citations
Related papers
- Transferable Perturbations of Deep Feature DistributionsNathan Inkawhich, Kevin J. Liang, Lawrence Carin, Yiran ChenICLR 2020 · 100 citations
- DVERGE: Diversifying Vulnerabilities for Enhanced Robust Generation of EnsemblesHuanrui Yang, Jingyang Zhang, Hongliang Dong, Nathan Inkawhich et al.NeurIPS 2020 · 144 citations
- Enhancing Cross-Task Black-Box Transferability of Adversarial Examples With Dispersion ReductionYantao Lu, Yunhan Jia, Jianyu Wang, Bai Li et al.CVPR 2020
- Perturbing Across the Feature Hierarchy to Improve Standard and Strict Blackbox Attack TransferabilityNathan Inkawhich, Kevin J. Liang, Binghui Wang, Matthew Inkawhich et al.NeurIPS 2020 · 105 citations
- Causes and Consequences of Representational Similarity in Machine Learning ModelsZeyu Michael Li, Hung Anh Vu, Damilola Awofisayo, Emily WengerICML 2026 · 1 citation
