On the Trade-off of Intra-/Inter-class Diversity for Supervised Pre-training
Jieyu Zhang, Bohan Wang, Zhengyu Hu, Pang Wei Koh, Alexander J. Ratner
Abstract
Pre-training datasets are critical for building state-of-the-art machine learning models, motivating rigorous study on their impact on downstream tasks. In this work, we study the impact of the trade-off between the intra-class diversity (the number of samples per class) and the inter-class diversity (the number of classes) of a supervised pre-training dataset. Empirically, we found that with the size of the pre-training dataset fixed, the best downstream performance comes with a balance on the intra-/inter-class diversity. To understand the underlying mechanism, we show theoretically that the downstream performance depends monotonically on both types of diversity. Notably, our theory reveals that the optimal class-to-sample ratio (#classes / #samples per class) is invariant to the size of the pre-training dataset, which motivates an application of predicting the optimal number of pre-training classes. We demonstrate the effectiveness of this application by an improvement of around 2 points on the downstream tasks when using ImageNet as the pre-training dataset.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext d8cd0c01-ab98-4b0f-9d54-9696472fe77dCited by top-tier papers2
- Understanding and Mitigating the Label Noise in Pre-training on Downstream TasksHao Chen, Jindong Wang, Ankit Shah, Ran Tao et al.ICLR 2024 · 49 citations
- On the Generalization Ability of Next-Token-Prediction PretrainingZhihao Li, Xue Jiang, Liyuan Liu, Xuelin Zhang et al.ICML 2025
Builds on7
- Learning Transferable Visual Models From Natural Language SupervisionAlec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh et al.ICML 2021 · 47,906 citations
- An Image is Worth 16x16 Words: Transformers for Image Recognition at ScaleAlexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn et al.ICLR 2021 · 21,477 citations
- Flamingo: a Visual Language Model for Few-Shot LearningJean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech et al.NeurIPS 2022 · 6,707 citations
- Moment Matching for Multi-Source Domain AdaptationXingchao Peng, Qinxun Bai, Xide Xia, Zijun Huang et al.ICCV 2019 · 2,239 citations
- Few-Shot Learning via Learning the Representation, ProvablySimon Shaolei Du, Wei Hu, Sham M. Kakade, Jason D. Lee et al.ICLR 2021 · 56 citations
Related papers
- On the Connection between Pre-training Data Diversity and Fine-tuning RobustnessVivek Ramanujan, Thao Nguyen, Sewoong Oh, Ali Farhadi et al.NeurIPS 2023 · 40 citations
- Exploring the Limits of Large Scale Pre-trainingSamira Abnar, Mostafa Dehghani, Behnam Neyshabur, Hanie SedghiICLR 2022 · 135 citations
- Theoretical Insights Into Multiclass Classification: A High-dimensional Asymptotic ViewChristos Thrampoulidis, Samet Oymak, Mahdi SoltanolkotabiNeurIPS 2020 · 46 citations
- Rethinking Supervised Pre-Training for Better Downstream TransferringYutong Feng, Jianwen Jiang, Mingqian Tang, Rong Jin et al.ICLR 2022 · 51 citations
- Fine-Tuning is Fine, if CalibratedZheda Mai, Arpita Chowdhury, Ping Zhang, Cheng-Hao Tu et al.NeurIPS 2024 · 34 citations
