Grounding inductive biases in natural images: invariance stems from variations in data
Diane Bouchacourt, Mark Ibrahim, Ari S. Morcos
Abstract
To perform well on unseen and potentially out-of-distribution samples, it is desirable for machine learning models to have a predictable response with respect to transformations affecting the factors of variation of the input. Here, we study the relative importance of several types of inductive biases towards such predictable behavior: the choice of data, their augmentations, and model architectures. Invariance is commonly achieved through hand-engineered data augmentation, but do standard data augmentations address transformations that explain variations in real data? While prior work has focused on synthetic data, we attempt here to characterize the factors of variation in a real dataset, ImageNet, and study the invariance of both standard residual networks and the recently proposed vision transformer with respect to changes in these factors. We show standard augmentation relies on a precise combination of translation and scale, with translation recapturing most of the performance improvement-despite the (approximate) translation invariance built in to convolutional architectures, such as residual networks. In fact, we found that scale and translation invariance was similar across residual networks and vision transformer models despite their markedly different architectural inductive biases. We show the training data itself is the main source of invariance, and that data augmentation only further increases the learned invariances. Notably, the invariances learned during training align with the ImageNet factors of variation we found. Finally, we find that the main factors of variation in ImageNet mostly relate to appearance and are specific to each class. * equal contribution, a coin was flipped 35th Conference on Neural Information Processing Systems (NeurIPS 2021),
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 5a65a492-df93-42cf-9910-b6f957b24d49Cited by top-tier papers13
- 3D molecule generation by denoising voxel gridsPedro O. Pinheiro, Joshua A. Rackers, Joseph Kleinhenz, Michael Maser et al.NeurIPS 2023 · 55 citations
- Understanding the detrimental class-level effects of data augmentationPolina Kirichenko, Mark Ibrahim, Randall Balestriero, Diane Bouchacourt et al.NeurIPS 2023 · 25 citations
- A Generative Model of Symmetry TransformationsJames Urquhart Allingham, Bruno Mlodozeniec, Shreyas Padhy, Javier Antorán et al.NeurIPS 2024 · 16 citations
- ImageNet-X: Understanding Model Mistakes with Factor of Variation AnnotationsBadr Youbi Idrissi, Diane Bouchacourt, Randall Balestriero, Ivan Evtimov et al.ICLR 2023 · 11 citations
- Deep invariant networks with differentiable augmentation layersCédric Rommel, Thomas Moreau, Alexandre GramfortNeurIPS 2022 · 11 citations
Builds on3
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Meta-learning Symmetries by ReparameterizationAllan Zhou, Tom Knowles, Chelsea FinnICLR 2021 · 105 citations
Related papers
- The Lie Derivative for Measuring Learned EquivarianceNate Gruver, Marc Anton Finzi, Micah Goldblum, Andrew Gordon WilsonICLR 2023 · 6 citations
- Understanding and Improving Robustness of Vision Transformers through Patch-based Negative AugmentationYao Qin, Chiyuan Zhang, Ting Chen, Balaji Lakshminarayanan et al.NeurIPS 2022 · 68 citations
- On the Strong Correlation Between Model Invariance and GeneralizationWeijian Deng, Stephen Gould, Liang ZhengNeurIPS 2022 · 28 citations
- Learning to Transform for Generalizable Instance-wise InvarianceUtkarsh Singhal, Carlos Esteves, Ameesh Makadia, Stella X. YuICCV 2023 · 3 citations
- Can CNNs Be More Robust Than Transformers?Zeyu Wang, Yutong Bai, Yuyin Zhou, Cihang XieICLR 2023 · 14 citations
