Grounding inductive biases in natural images: invariance stems from variations in data
Diane Bouchacourt, Mark Ibrahim, Ari S. Morcos
摘要
To perform well on unseen and potentially out-of-distribution samples, it is desirable for machine learning models to have a predictable response with respect to transformations affecting the factors of variation of the input. Here, we study the relative importance of several types of inductive biases towards such predictable behavior: the choice of data, their augmentations, and model architectures. Invariance is commonly achieved through hand-engineered data augmentation, but do standard data augmentations address transformations that explain variations in real data? While prior work has focused on synthetic data, we attempt here to characterize the factors of variation in a real dataset, ImageNet, and study the invariance of both standard residual networks and the recently proposed vision transformer with respect to changes in these factors. We show standard augmentation relies on a precise combination of translation and scale, with translation recapturing most of the performance improvement-despite the (approximate) translation invariance built in to convolutional architectures, such as residual networks. In fact, we found that scale and translation invariance was similar across residual networks and vision transformer models despite their markedly different architectural inductive biases. We show the training data itself is the main source of invariance, and that data augmentation only further increases the learned invariances. Notably, the invariances learned during training align with the ImageNet factors of variation we found. Finally, we find that the main factors of variation in ImageNet mostly relate to appearance and are specific to each class. * equal contribution, a coin was flipped 35th Conference on Neural Information Processing Systems (NeurIPS 2021),
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper13
- 3D molecule generation by denoising voxel gridsPedro O. Pinheiro, Joshua A. Rackers, Joseph Kleinhenz, Michael Maser 等NeurIPS 2023 · 被引用 55 次
- Understanding the detrimental class-level effects of data augmentationPolina Kirichenko, Mark Ibrahim, Randall Balestriero, Diane Bouchacourt 等NeurIPS 2023 · 被引用 25 次
- A Generative Model of Symmetry TransformationsJames Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán 等NeurIPS 2024 · 被引用 16 次
- ImageNet-X: Understanding Model Mistakes with Factor of Variation AnnotationsBadr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov 等ICLR 2023 · 被引用 11 次
- Deep invariant networks with differentiable augmentation layersCédric Rommel, Thomas Moreau, Alexandre GramfortNeurIPS 2022 · 被引用 11 次
它引用的顶会 Paper3
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn 等ICLR 2021 · 被引用 21,477 次
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal 等NeurIPS 2020 · 被引用 5,249 次
- Meta-learning Symmetries by ReparameterizationAllan Zhou, Tom Knowles, Chelsea FinnICLR 2021 · 被引用 105 次
相关 Paper
- The Lie Derivative for Measuring Learned EquivarianceNate Gruver, Marc Anton Finzi, Micah Goldblum, Andrew Gordon WilsonICLR 2023 · 被引用 6 次
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative AugmentationYao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan 等NeurIPS 2022 · 被引用 68 次
- On the Strong Correlation Between Model Invariance and GeneralizationWeijian Deng, Stephen Gould, Liang ZhengNeurIPS 2022 · 被引用 28 次
- Learning to Transform for Generalizable Instance-wise InvarianceUtkarsh Singhal, Carlos Esteves, Ameesh Makadia, Stella X. YuICCV 2023 · 被引用 3 次
- Can CNNs Be More Robust Than Transformers?Zeyu Wang, Yutong Bai, Yuyin Zhou, Cihang XieICLR 2023 · 被引用 14 次
