No One Representation to Rule Them All: Overlapping Features of Training Methods
Raphael Gontijo Lopes, Yann N. Dauphin, Ekin Dogus Cubuk
Abstract
Despite being able to capture a range of features of the data, high accuracy models trained with supervision tend to make similar predictions. This seemingly implies that high-performing models share similar biases regardless of training methodology, which would limit ensembling benefits and render low-accuracy models as having little practical use. Against this backdrop, recent work has developed quite different training techniques, such as large-scale contrastive learning, yielding competitively high accuracy on generalization and robustness benchmarks. This motivates us to revisit the assumption that models necessarily learn similar functions. We conduct a large-scale empirical study of models across hyper-parameters, architectures, frameworks, and datasets. We find that model pairs that diverge more in training methodology display categorically different generalization behavior, producing increasingly uncorrelated errors. We show these models specialize in subdomains of the data, leading to higher ensemble performance: with just 2 models (each with ImageNet accuracy 76.5%), we can create ensembles with 83.4% (+7% boost). Surprisingly, we find that even significantly low-accuracy models can be used to improve high-accuracy models. Finally, we show diverging training methodology yield representations that capture overlapping (but not supersetting) feature sets which, when combined, lead to increased downstream performance. How does training methodology affect learned representations and prediction behavior?
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers29
- Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference timeMitchell Wortsman, Gabriel Ilharco, Samir Yitzhak Gadre, Rebecca Roelofs et al.ICML 2022 · 1,464 citations
- Robust fine-tuning of zero-shot modelsMitchell Wortsman, Gabriel Ilharco, Jong Wook Kim, Mike Li et al.CVPR 2022 · 364 citations
- Diverse Weight Averaging for Out-of-Distribution GeneralizationAlexandre Ramé, Matthieu Kirchmeyer, Thibaud Rahier, Alain Rakotomamonjy et al.NeurIPS 2022 · 183 citations
- WARM: On the Benefits of Weight Averaged Reward ModelsAlexandre Ramé, Nino Vieillard, Léonard Hussenot, Robert Dadashi et al.ICML 2024 · 145 citations
- Quality Not Quantity: On the Interaction between Dataset Design and Robustness of CLIPThao Nguyen, Gabriel Ilharco, Mitchell Wortsman, Sewoong Oh et al.NeurIPS 2022 · 131 citations
Builds on14
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Training data-efficient image transformers & distillation through attentionHugo Touvron, Matthieu Cord, Matthijs Douze, Francisco Massa et al.ICML 2021 · 8,974 citations
- Scaling Up Visual and Vision-Language Representation Learning With Noisy Text SupervisionChao Jia, Yinfei Yang, Ye Xia, Yi-Ting Chen et al.ICML 2021 · 5,401 citations
Related papers
- Multi-Symmetry Ensembles: Improving Diversity and Generalization via Opposing SymmetriesCharlotte Loh, Seungwook Han, Shivchander Sudalairaj, Rumen Dangovski et al.ICML 2023 · 2 citations
- Ensembles of Locally Independent Prediction ModelsAndrew Slavin Ross, Weiwei Pan, Leo A. Celi, Finale Doshi-VelezAAAI 2020 · 33 citations
- Ex Uno Pluria: Insights on Ensembling in Low Precision Number SystemsGiung Nam, Juho LeeNeurIPS 2024 · 2 citations
- The Trade-off between Universality and Label Efficiency of Representations from Contrastive LearningZhenmei Shi, Jiefeng Chen, Kunyang Li, Jayaram Raghuram et al.ICLR 2023 · 1 citation
- Synergy and Diversity in CLIP: Enhancing Performance Through Adaptive Backbone EnsemblingCristian Rodriguez Opazo, Ehsan Abbasnejad, Damien Teney, Hamed Damirchi et al.ICLR 2025
