Understanding Invariance via Feedforward Inversion of Discriminatively Trained Classifiers
Piotr Teterwak, Chiyuan Zhang, Dilip Krishnan, Michael C. Mozer
Abstract
A discriminatively trained neural net classifier can fit the training data perfectly if all information about its input other than class membership has been discarded prior to the output layer. Surprisingly, past research has discovered that some extraneous visual detail remains in the logit vector. This finding is based on inversion techniques that map deep embeddings back to images. We explore this phenomenon further using a novel synthesis of methods, yielding a feedforward inversion model that produces remarkably high fidelity reconstructions, qualitatively superior to those of past efforts. When applied to an adversarially robust classifier model, the reconstructions contain sufficient local detail and global structure that they might be confused with the original image in a quick glance, and the object category can clearly be gleaned from the reconstruction. Our approach is based on Big-GAN (Brock, 2019), with conditioning on logits instead of one-hot class labels. We use our reconstruction model as a tool for exploring the nature of representations, including: the influence of model architecture and training objectives (specifically robust losses), the forms of invariance that networks achieve, representational differences between correctly and incorrectly classified images, and the effects of manipulating logits and images. We believe that our method can inspire future investigations into the nature of information flow in a neural net and can provide diagnostics for improving discriminative models. We provide pre-trained models and visualizations at https://sites.google.com/view/ understanding-invariance/home .
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers6
- Deduplicating Training Data Makes Language Models BetterKatherine Lee, Daphne Ippolito, Andrew Nystrom, Chiyuan Zhang et al.ACL 2022 · 844 citations
- Can Neural Network Memorization Be Localized?Pratyush Maini, Michael Curtis Mozer, Hanie Sedghi, Zachary Chase Lipton et al.ICML 2023 · 82 citations
- Spectral Bias in Practice: The Role of Function Frequency in GeneralizationSara Fridovich-Keil, Raphael Gontijo Lopes, Rebecca RoelofsNeurIPS 2022 · 61 citations
- Better Language Model Inversion by Compactly Representing Next-Token DistributionsMurtaza Nazir, Matthew Finlayson, John X. Morris, Xiang Ren et al.NeurIPS 2025 · 13 citations
- Language Model InversionJohn X. Morris, Wenting Zhao, Justin T. Chiu, Vitaly Shmatikov et al.ICLR 2024 · 6 citations
Builds on5
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer et al.NeurIPS 2021 · 3,862 citations
- Boundless: Generative Adversarial Networks for Image ExtensionDilip Krishnan, Piotr Teterwak, Aaron Sarna, Aaron Maschinot et al.ICCV 2019 · 129 citations
- von Mises-Fisher Loss: An Exploration of Embedding Geometries for Supervised LearningTyler R. Scott, Andrew C. Gallagher, Michael C. MozerICCV 2021 · 56 citations
- Semantic Pyramid for Image GenerationAssaf Shocher, Yossi Gandelsman, Inbar Mosseri, Michal Yarom et al.CVPR 2020
Related papers
- Adversarial Training Reduces Information and Improves TransferabilityMatteo Terzi, Alessandro Achille, Marco Maggipinto, Gian Antonio SustoAAAI 2021 · 25 citations
- On the Functional Similarity of Robust and Non-Robust Neural RepresentationsAndrás Balogh, Márk JelasityICML 2023 · 4 citations
- Neural Representations Reveal Distinct Modes of Class Fitting in Residual Convolutional NetworksMichal Jamroz, Marcin KurdzielAAAI 2023
- Attack to Explain Deep RepresentationMohammad A. A. K. Jalwana, Naveed Akhtar, Mohammed Bennamoun, Ajmal MianCVPR 2020
- Adversarial Attacks are Reversible with Natural SupervisionChengzhi Mao, Mia Chiquier, Hao Wang, Junfeng Yang et al.ICCV 2021 · 66 citations
