Neural Representations Reveal Distinct Modes of Class Fitting in Residual Convolutional Networks
Michal Jamroz, Marcin Kurdziel
摘要
We leverage probabilistic models of neural representations to investigate how residual networks fit classes. To this end, we estimate class-conditional density models for representations learned by deep ResNets. We then use these models to characterize distributions of representations across learned classes. Surprisingly, we find that classes in the investigated models are not fitted in an uniform way. On the contrary: we uncover two groups of classes that are fitted with markedly different distributions of representations. These distinct modes of class-fitting are evident only in the deeper layers of the investigated models, indicating that they are not related to low-level image features. We show that the uncovered structure in neural representations correlate with memorization of training examples and adversarial robustness. Finally, we compare class-conditional distributions of neural representations between memorized and typical examples. This allows us to uncover where in the network structure class labels arise for memorized and standard inputs.
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
它引用的顶会 Paper5
- MLP-Mixer: An all-MLP Architecture for VisionIlya O. Tolstikhin, Neil Houlsby, Alexander Kolesnikov, Lucas Beyer 等NeurIPS 2021 · 被引用 3,862 次
- Sharpness-aware Minimization for Efficiently Improving GeneralizationPierre Foret, Ariel Kleiner, Hossein Mobahi, Behnam NeyshaburICLR 2021 · 被引用 1,861 次
- What Neural Networks Memorize and Why: Discovering the Long Tail via Influence EstimationVitaly Feldman, Chiyuan ZhangNeurIPS 2020 · 被引用 674 次
- Does learning require memorization? a short tale about a long tailVitaly FeldmanSTOC 2020 · 被引用 28 次
- A Bayesian Nonparametrics View into Deep RepresentationsMichal Jamroz, Marcin Kurdziel, Mateusz OpalaNeurIPS 2020 · 被引用 1 次
相关 Paper
- On the Functional Similarity of Robust and Non-Robust Neural RepresentationsAndrás Balogh, Márk JelasityICML 2023 · 被引用 4 次
- Identifying and Understanding Cross-Class Features in Adversarial TrainingZeming Wei, Steven Y. Guo, Yisen WangICML 2025
- Understanding Invariance via Feedforward Inversion of Discriminatively Trained ClassifiersPiotr Teterwak, Chiyuan Zhang, Dilip Krishnan, Michael C. MozerICML 2021 · 被引用 11 次
- On the Clean Generalization and Robust Overfitting in Adversarial Training from Two Theoretical Views: Representation Complexity and Training DynamicsBinghui Li, Yuanzhi LiICML 2025
- Understanding Robust Learning through the Lens of Representation SimilaritiesChristian Cianfarani, Arjun Nitin Bhagoji, Vikash Sehwag, Ben Y. Zhao 等NeurIPS 2022 · 被引用 20 次
