Meta-learning to Improve Pre-training
Aniruddh Raghu, Jonathan Lorraine, Simon Kornblith, Matthew McDermott, David Duvenaud
摘要
Pre-training (PT) followed by fine-tuning (FT) is an effective method for training neural networks, and has led to significant performance improvements in many domains. PT can incorporate various design choices such as task and data reweighting strategies, augmentation policies, and noise models, all of which can significantly impact the quality of representations learned. The hyperparameters introduced by these strategies therefore must be tuned appropriately. However, setting the values of these hyperparameters is challenging. Most existing methods either struggle to scale to high dimensions, are too slow and memory-intensive, or cannot be directly applied to the two-stage PT and FT learning process. In this work, we propose an efficient, gradient-based algorithm to meta-learn PT hyperparameters. We formalize the PT hyperparameter optimization problem and propose a novel method to obtain PT hyperparameter gradients by combining implicit differentiation and backpropagation through unrolled optimization. We demonstrate that our method improves predictive performance on two real-world domains. First, we optimize high-dimensional task weighting hyperparameters for multitask pre-training on protein-protein interaction graphs and improve AUROC by up to 3.9%. Second, we optimize a data augmentation neural network for self-supervised PT with SimCLR on electrocardiography data and improve AUROC by up to 1.9%. Choosing optimal PT hyperparameter values is challenging, and existing methods do not work well. Simple approaches such as random or grid search are inefficient since evaluating a hyperparameter setting requires performing the full, two-stage PT & FT optimization, which may be prohibitively computationally expensive. Gradient-free approaches, such as Bayesian optimization or evolutionary algorithms [33, 61, 47] , are also limited in how well they scale to this setting. Gradient-based 35th Conference on Neural Information Processing Systems (NeurIPS 2021).
问问这篇 Paper
智能体会读完全文。
Lune 把这篇 Paper 索引到了每一个公式,引用它的顶会 Paper 也一样。你提问,回答直接引用原文。
引用它的顶会 Paper11
- Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language ModelsHong Liu, Sang Michael Xie, Zhiyuan Li, Tengyu MaICML 2023 · 被引用 82 次
- Sequential Multi-Dimensional Self-Supervised Learning for Clinical Time SeriesAniruddh Raghu, Payal Chandak, Ridwan Alam, John V. Guttag 等ICML 2023 · 被引用 18 次
- Task Discovery: Finding the Tasks that Neural Networks Generalize onAndrei Atanov, Andrei Filatov, Teresa Yeo, Ajay Sohmshetty 等NeurIPS 2022 · 被引用 11 次
- Joint Attribute and Model Generalization Learning for Privacy-Preserving Action RecognitionDuo Peng, Li Xu, Qiuhong Ke, Ping Hu 等NeurIPS 2023 · 被引用 8 次
- Efficient Event Camera Data Pretraining with Adaptive Prompt FusionQuanmin Liang, Qiang Li, Shuai Liu, Xinzi Cao 等ICCV 2025 · 被引用 6 次
它引用的顶会 Paper16
- A Simple Framework for Contrastive Learning of Visual RepresentationsTing Chen, Simon Kornblith, Mohammad Norouzi, Geoffrey E. HintonICML 2020 · 被引用 24,064 次
- Big Self-Supervised Models are Strong Semi-Supervised LearnersTing Chen, Simon Kornblith, Kevin Swersky, Mohammad Norouzi 等NeurIPS 2020 · 被引用 2,611 次
- Gradient Surgery for Multi-Task LearningTianhe Yu, Saurabh Kumar, Abhishek Gupta, Sergey Levine 等NeurIPS 2020 · 被引用 2,261 次
- Strategies for Pre-training Graph Neural NetworksWeihua Hu, Bowen Liu, Joseph Gomes, Marinka Zitnik 等ICLR 2020 · 被引用 1,744 次
- Data-Efficient Image Recognition with Contrastive Predictive CodingOlivier J. HénaffICML 2020 · 被引用 1,553 次
相关 Paper
- Quick-Tune: Quickly Learning Which Pretrained Model to Finetune and HowSebastian Pineda-Arango, Fabio Ferreira, Arlind Kadra, Frank Hutter 等ICLR 2024 · 被引用 27 次
- Meta Fine-Tuning Neural Language Models for Multi-Domain Text MiningChengyu Wang, Minghui Qiu, Jun Huang, Xiaofeng HeEMNLP 2020 · 被引用 19 次
- Learning to Initialize: Can Meta Learning Improve Cross-task Generalization in Prompt Tuning?Chengwei Qin, Shafiq R. Joty, Qian Li, Ruochen ZhaoACL 2023 · 被引用 8 次
- LiFT: Learning to Fine-Tune via Bayesian Parameter Efficient Meta Fine-TuningMinyoung Kim, Timothy M. HospedalesICLR 2025
- Learning to Select Best Forecast Tasks for Clinical Outcome PredictionYuan Xue, Nan Du, Anne Mottram, Martin Seneviratne 等NeurIPS 2020 · 被引用 9 次
