Do Invariances in Deep Neural Networks Align with Human Perception?
Vedant Nanda, Ayan Majumdar, Camila Kolling, John P. Dickerson, Krishna P. Gummadi, Bradley C. Love, Adrian Weller
摘要
An evaluation criterion for safe and trustworthy deep learning is how well the invariances captured by representations of deep neural networks (DNNs) are shared with humans. We identify challenges in measuring these invariances. Prior works used gradient-based methods to generate identically represented inputs (IRIs), ie, inputs which have identical representations (on a given layer) of a neural network, and thus capture invariances of a given network. One necessary criterion for a network's invariances to align with human perception is for its IRIs look 'similar' to humans. Prior works, however, have mixed takeaways; some argue that later layers of DNNs do not learn human-like invariances yet others seem to indicate otherwise. We argue that the loss function used to generate IRIs can heavily affect takeaways about invariances of the network and is the primary reason for these conflicting findings. We propose an adversarial regularizer on the IRI generation loss that finds IRIs that make any model appear to have very little shared invariance with humans. Based on this evidence, we argue that there is scope for improving models to have human-like invariances, and further, to have meaningful comparisons between models one should use IRIs generated using the regularizer-free loss. We then conduct an in-depth investigation of how different components (eg architectures, training losses, data augmentations) of the deep learning pipeline contribute to learning models that have good alignment with humans. We find that architectures with residual connections trained using a (self-supervised) contrastive loss with l_p ball adversarial data augmentation tend to learn invariances that are most aligned with humans. Code: github.com/nvedant07/Human-NN-Alignment
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper2
- Distribution Shift Is Key to Learning Invariant PredictionHong Zheng, Fei TengAAAI 2026
- Synthesizing Images on Perceptual Boundaries of ANNs for Uncovering and Manipulating Human Perceptual VariabilityChen Wei, Chi Zhang, Jiachen Zou, Haotian Deng 等ICML 2025
它引用的顶会 Paper10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Distillation as a Defense to Adversarial Perturbations Against Deep Neural NetworksNicolas Papernot, Patrick D. McDaniel, Xi Wu, Somesh Jha 等S&P 2016 · 被引用 3,275 次
- Improving Adversarial Robustness Requires Revisiting Misclassified ExamplesYisen Wang, Difan Zou, Jinfeng Yi, James Bailey 等ICLR 2020 · 被引用 829 次
- Measuring Robustness to Natural Distribution Shifts in Image ClassificationRohan Taori, Achal Dave, Vaishaal Shankar, Nicholas Carlini 等NeurIPS 2020 · 被引用 731 次
- Do Adversarially Robust ImageNet Models Transfer Better?Hadi Salman, Andrew Ilyas, Logan Engstrom, Ashish Kapoor 等NeurIPS 2020 · 被引用 506 次
相关 Paper
- Measuring Representational Robustness of Neural Networks Through Shared InvariancesVedant Nanda, Till Speicher, Camila Kolling, John P. Dickerson 等ICML 2022 · 被引用 15 次
- Improving Transformation Invariance in Contrastive Representation LearningAdam Foster, Rattana Pukdee, Tom RainforthICLR 2021 · 被引用 25 次
- Enhanced Regularizers for Attributional RobustnessAnindya Sarkar, Anirban Sarkar, Vineeth N. BalasubramanianAAAI 2021 · 被引用 18 次
- Improving Equivariance in State-of-the-Art Supervised Depth and Normal PredictorsYuanyi Zhong, Anand Bhattad, Yu-Xiong Wang, David A. ForsythICCV 2023 · 被引用 3 次
- Low-Pass Filtering Improves Behavioral Alignment of Vision ModelsMax Wolff, Thomas Klein, Evgenia Rusak, Felix A. Wichmann 等ICLR 2026 · 被引用 2 次
