On the Provable Advantage of Unsupervised Pretraining
Jiawei Ge, Shange Tang, Jianqing Fan, Chi Jin
Abstract
Unsupervised pretraining, which learns a useful representation using a large amount of unlabeled data to facilitate the learning of downstream tasks, is a critical component of modern large-scale machine learning systems. Despite its tremendous empirical success, the rigorous theoretical understanding of why unsupervised pretraining generally helps remains rather limited -- most existing results are restricted to particular methods or approaches for unsupervised pretraining with specialized structural assumptions. This paper studies a generic framework, where the unsupervised representation learning task is specified by an abstract class of latent variable models and the downstream task is specified by a class of prediction functions . We consider a natural approach of using Maximum Likelihood Estimation (MLE) for unsupervised pretraining and Empirical Risk Minimization (ERM) for learning downstream tasks. We prove that, under a mild ''informative'' condition, our algorithm achieves an excess risk of for downstream tasks, where are complexity measures of function classes , and are the number of unlabeled and labeled data respectively. Comparing to the baseline of achieved by performing supervised learning using only the labeled data, our result rigorously shows the benefit of unsupervised pretraining when and . This paper further shows that our generic framework covers a wide range of approaches for unsupervised pretraining, including factor models, Gaussian mixture models, and contrastive learning.
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Cited by top-tier papers3
- A Theory of Multimodal LearningZhou LuNeurIPS 2023 · 48 citations
- Benign Overfitting in Out-of-Distribution Generalization of Linear ModelsShange Tang, Jiayun Wu, Jianqing Fan, Chi JinICLR 2025
- A Theory for Conditional Generative Modeling on Multiple Data SourcesRongzhen Wang, Yan Zhang, Chenyu Zheng, Chongxuan Li et al.ICML 2025
Builds on12
- wav2vec 2.0: A Framework for Self-Supervised Learning of Speech RepresentationsAlexei Baevski, Yuhao Zhou, Abdelrahman Mohamed, Michael AuliNeurIPS 2020 · 9,451 citations
- Provable Guarantees for Self-Supervised Deep Learning with Spectral Contrastive LossJeff Z. HaoChen, Colin Wei, Adrien Gaidon, Tengyu MaNeurIPS 2021 · 425 citations
- FLAMBE: Structural Complexity and Representation Learning of Low Rank MDPsAlekh Agarwal, Sham M. Kakade, Akshay Krishnamurthy, Wen SunNeurIPS 2020 · 271 citations
- On the Theory of Transfer Learning: The Importance of Task DiversityNilesh Tripuraneni, Michael I. Jordan, Chi JinNeurIPS 2020 · 263 citations
- Unsupervised Pre-Training of Image Features on Non-Curated DataMathilde Caron, Piotr Bojanowski, Julien Mairal, Armand JoulinICCV 2019 · 254 citations
Related papers
- The Trade-off between Universality and Label Efficiency of Representations from Contrastive LearningZhenmei Shi, Jiefeng Chen, Kunyang Li, Jayaram Raghuram et al.ICLR 2023 · 1 citation
- The Breakdown of Gaussian Universality in Classification of High-dimensional Linear Factor MixturesXiaoyi Mai, Zhenyu LiaoICLR 2025
- A Generalization Theory for Zero-Shot PredictionRonak Mehta, Zaïd HarchaouiICML 2025
- Towards Unsupervised Domain GeneralizationXingxuan Zhang, Linjun Zhou, Renzhe Xu, Peng Cui et al.CVPR 2022 · 43 citations
- Transformers are Minimax Optimal Nonparametric In-Context LearnersJuno Kim, Tai Nakamaki, Taiji SuzukiNeurIPS 2024 · 42 citations
