Meta-learning to Improve Pre-training
Aniruddh Raghu, Jonathan Lorraine, Simon Kornblith, Matthew McDermott, David Duvenaud
Abstract
Pre-training (PT) followed by fine-tuning (FT) is an effective method for training neural networks, and has led to significant performance improvements in many domains. PT can incorporate various design choices such as task and data reweighting strategies, augmentation policies, and noise models, all of which can significantly impact the quality of representations learned. The hyperparameters introduced by these strategies therefore must be tuned appropriately. However, setting the values of these hyperparameters is challenging. Most existing methods either struggle to scale to high dimensions, are too slow and memory-intensive, or cannot be directly applied to the two-stage PT and FT learning process. In this work, we propose an efficient, gradient-based algorithm to meta-learn PT hyperparameters. We formalize the PT hyperparameter optimization problem and propose a novel method to obtain PT hyperparameter gradients by combining implicit differentiation and backpropagation through unrolled optimization. We demonstrate that our method improves predictive performance on two real-world domains. First, we optimize high-dimensional task weighting hyperparameters for multitask pre-training on protein-protein interaction graphs and improve AUROC by up to 3.9%. Second, we optimize a data augmentation neural network for self-supervised PT with SimCLR on electrocardiography data and improve AUROC by up to 1.9%. Choosing optimal PT hyperparameter values is challenging, and existing methods do not work well. Simple approaches such as random or grid search are inefficient since evaluating a hyperparameter setting requires performing the full, two-stage PT & FT optimization, which may be prohibitively computationally expensive. Gradient-free approaches, such as Bayesian optimization or evolutionary algorithms [33, 61, 47] , are also limited in how well they scale to this setting. Gradient-based 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
Ask about this paper
Your agent reads all of it.
Lune indexed this paper to the last equation, along with the top-tier papers that cite it. Ask a question and the answer quotes them.
Your agent calls
Luneget_paper_fulltext
Free to start. No credit card required.
Terminal
Install the CLIlune papers fulltext 0e85880b-4893-4f78-92a1-3c67400d65feCited by top-tier papers11
- Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language ModelsHong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu MaICML 2023 · 82 citations
- Sequential Multi-Dimensional Self-Supervised Learning for Clinical Time SeriesAniruddh Raghu, Payal Chandak, Ridwan Alam, John V. Guttag et al.ICML 2023 · 18 citations
- Task Discovery: Finding the Tasks that Neural Networks Generalize onAndrei Atanov, Andrei Filatov, Teresa Yeo, Ajay Sohmshetty et al.NeurIPS 2022 · 11 citations
- Joint Attribute and Model Generalization Learning for Privacy-Preserving Action RecognitionDuo Peng, Li Xu, Qiuhong Ke, Ping Hu et al.NeurIPS 2023 · 8 citations
- Efficient Event Camera Data Pretraining with Adaptive Prompt FusionQuanmin Liang, Qiang Li, Shuai Liu, Xinzi Cao et al.ICCV 2025 · 6 citations
Builds on16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 24,064 citations
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi et al.NeurIPS 2020 · 2,611 citations
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine et al.NeurIPS 2020 · 2,261 citations
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik et al.ICLR 2020 · 1,744 citations
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 1,553 citations
Related papers
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter et al.ICLR 2024 · 27 citations
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 19 citations
- Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?Chengwei Qin, Shafiq R. Joty, Qian Li, Ruochen ZhaoACL 2023 · 8 citations
- LiFT: Learning to Fine-Tune via Bayesian Parameter Efficient Meta Fine-TuningMinyoung Kim, Timothy M. HospedalesICLR 2025
- Learning to Select Best Forecast Tasks for Clinical Outcome PredictionYuan Xue, Nan Du, Anne Mottram, Martin Seneviratne et al.NeurIPS 2020 · 9 citations
