Adversarial Training Reduces Information and Improves Transferability
Matteo Terzi, Alessandro Achille, Marco Maggipinto, Gian Antonio Susto
Abstract
Recent results show that features of adversarially trained networks for classification, in addition to being robust, enable desirable properties such as invertibility. The latter property may seem counter-intuitive as it is widely accepted by the community that classification models should only capture the minimal information (features) required for the task. Motivated by this discrepancy, we investigate the dual relationship between Adversarial Training and Information Theory. We show that the Adversarial Training can improve linear transferability to new tasks, from which arises a new trade-off between transferability of representations and accuracy on the source task. We validate our results employing robust networks trained on CIFAR-10, CIFAR-100 and ImageNet on several datasets. Moreover, we show that Adversarial Training reduces Fisher information of representations about the input and of the weights about the task, and we provide a theoretical argument which explains the invertibility of deterministic networks without violating the principle of minimality. Finally, we leverage our theoretical insights to remarkably improve the quality of reconstructed images through inversion.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 2c9682e5-3e5f-414c-912e-90b3ba483c5aCited by top-tier papers7
- Better Safe Than Sorry: Preventing Delusive Adversaries with Adversarial TrainingLue Tao, Lei Feng, Jinfeng Yi, Sheng-Jun Huang et al.NeurIPS 2021 · 90 citations
- A Little Robustness Goes a Long Way: Leveraging Robust Features for Targeted Transfer AttacksJacob M. Springer, Melanie Mitchell, Garrett T. KenyonNeurIPS 2021 · 54 citations
- Batch Normalization Increases Adversarial Vulnerability and Decreases Adversarial Transferability: A Non-Robust Feature PerspectivePhilipp Benz, Chaoning Zhang, In So KweonICCV 2021 · 47 citations
- MAGIC: Mask-Guided Image Synthesis by Inverting a Quasi-robust ClassifierMozhdeh Rouhsedaghat, Masoud Monajatipoor, C.-C. Jay Kuo, Iacopo MasiAAAI 2023 · 9 citations
- On the Functional Similarity of Robust and Non-Robust Neural RepresentationsAndrás Balogh, Márk JelasityICML 2023 · 4 citations
Builds on1
Related papers
- Adversarially-Trained Deep Nets Transfer Better: Illustration on Image ClassificationFrancisco Utrera, Evan Kravitz, N. Benjamin Erichson, Rajiv Khanna et al.ICLR 2021 · 42 citations
- Learning Adversarially Robust Representations via Worst-Case Mutual Information MaximizationSicheng Zhu, Xiao Zhang, David EvansICML 2020 · 30 citations
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor et al.NeurIPS 2020 · 506 citations
- Understanding Invariance via Feedforward Inversion of Discriminatively Trained ClassifiersPiotr Teterwak, Chiyuan Zhang, Dilip Krishnan, Michael C. MozerICML 2021 · 11 citations
- Reducing information dependency does not cause training data privacy. Adversarially non-robust features do.Rasmus Torp, Shailen Smith, Adam BreuerICLR 2026 · 1 citation
