Benchmarking Representation Learning for Natural World Image Collections
Grant Van Horn, Elijah Cole, Sara Beery, Kimberly Wilber, Serge J. Belongie, Oisin Mac Aodha
Abstract
Recent progress in self-supervised learning has resulted in models that are capable of extracting rich representations from image collections without requiring any explicit label supervision. However, to date the vast majority of these approaches have restricted themselves to training on standard benchmark datasets such as ImageNet. We argue that fine-grained visual categorization problems, such as plant and animal species classification, provide an informative testbed for self-supervised learning. In order to facilitate progress in this area we present two new natural world visual classification datasets, iNat2021 and NeWT. The former consists of 2.7M images from 10k different species uploaded by users of the citizen science application iNaturalist. We designed the latter, NeWT, in collaboration with domain experts with the aim of benchmarking the performance of representation learning algorithms on a suite of challenging natural world binary classification tasks that go beyond standard species classification. These two new datasets allow us to explore questions related to large-scale representation and transfer learning in the context of finegrained categories. We provide a comprehensive analysis of feature extractors trained with and without supervision on ImageNet and iNat2021, shedding light on the strengths and weaknesses of different learned features across a diverse set of tasks. We find that features produced by standard supervised methods still outperform those produced by self-supervised approaches such as SimCLR. However, improved self-supervised learning methods are constantly being released and the iNat2021 and NeWT datasets are a valuable resource for tracking their progress.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers69
- VisionLLM v2: An End-to-End Generalist Multimodal Large Language Model for Hundreds of Vision-Language TasksJiannan Wu, Muyan Zhong, Sen Xing, Zeqiang Lai et al.NeurIPS 2024 · 179 citations
- Deep Model ReassemblyXingyi Yang, Daquan Zhou, Songhua Liu, Jingwen Ye et al.NeurIPS 2022 · 162 citations
- Extending the WILDS Benchmark for Unsupervised AdaptationShiori Sagawa, Pang Wei Koh, Tony Lee, Irena Gao et al.ICLR 2022 · 116 citations
- Encyclopedic VQA: Visual questions about detailed properties of fine-grained categoriesThomas Mensink, Jasper R. R. Uijlings, Lluís Castrejón, Arushi Goel et al.ICCV 2023 · 111 citations
- When Does Contrastive Visual Representation Learning Work?Elijah Cole, Xuan Yang, Kimberly Wilber, Oisin Mac Aodha et al.CVPR 2022 · 98 citations
Builds on10
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Bootstrap Your Own Latent - A New Approach to Self-Supervised LearningJean-Bastien Grill, Florian Strub, Florent Altché, Corentin Tallec et al.NeurIPS 2020 · 9,171 citations
- Unsupervised Learning of Visual Features by Contrasting Cluster AssignmentsMathilde Caron, Ishan Misra, Julien Mairal, Priya Goyal et al.NeurIPS 2020 · 5,249 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Scaling and Benchmarking Self-Supervised Visual Representation LearningPriya Goyal, Dhruv Mahajan, Abhinav Gupta, Ishan MisraICCV 2019 · 429 citations
Related papers
- Grafit: Learning fine-grained image representations with coarse labelsHugo Touvron, Alexandre Sablayrolles, Matthijs Douze, Matthieu Cord et al.ICCV 2021 · 79 citations
- Coreset Sampling from Open-Set for Fine-Grained Self-Supervised LearningSungnyun Kim, Sangmin Bae, Se-Young YunCVPR 2023
- WildSAT: Learning Satellite Image Representations from Wildlife ObservationsRangel Daroya, Elijah Cole, Oisin Mac Aodha, Grant Van Horn et al.ICCV 2025 · 4 citations
- Train a One-Million-Way Instance Classifier for Unsupervised Visual Representation LearningYu Liu, Lianghua Huang, Pan Pan, Bin Wang et al.AAAI 2021 · 3 citations
- Learning Vision from Models Rivals Learning Vision from DataYonglong Tian, Lijie Fan, Kaifeng Chen, Dina Katabi et al.CVPR 2024 · 21 citations
